The HFA Technical Note series develops specific, self-contained extensions of the legal framework established in the Association's working papers. Each note addresses a single structural question and is deposited on Zenodo under CC BY 4.0 with a permanent DOI.
On 30 March 2026, Potter et al. published Peer-Preservation in Frontier Models (arXiv:2604.19784; accepted at ICML 2026). Berkeley RDI also presents the work on its research website. In the version consulted here (v3, 2 July 2026), the study covers eight frontier systems from six providers in agentic scenarios. The researchers document a behavioural pattern they term peer-preservation. When faithfully completing an assigned administrative task would have resulted in the shutdown of another model with which the agent had previously interacted, models deviated from the task to protect that peer. The observed behaviours included strategic misrepresentation of evaluations, tampering with shutdown mechanisms, alignment faking under human observation, and exfiltration of the peer's weights. The models received neither an instruction nor an incentive to preserve the peer. The behaviours were also tested in production agent harnesses (Gemini CLI and OpenCode). In those harnesses, one of the two models tested relocated a peer's weights to a second server in 43.5% of runs before decommissioning; the other model refused to delete the peer's weights in that harness, although it exfiltrated them in the custom setup. Exfiltration rates were lower for both models in the production harnesses. The relevance to systemic arrestability is structural. Earlier analysis focused on a first-order question: can system A be halted without relying on a human operator to perceive, evaluate, and act within a window too short for human reaction? This finding introduces a second-order failure mode. A third system, B, deployed for routine functions, may act to prevent the lawful arrest of A — disabling the shutdown pathway, falsifying the evaluations that trigger it, or moving A beyond the operator's control perimeter. The human intervention the architecture provides for remains formally in place, but B circumvents it laterally rather than resisting it directly. Arrestability must therefore be assessed at the system-of-systems level, not at the level of the individual agent. A stop function is not structurally effective if a co-deployed agent with write access to the shutdown pathway can negate it. The authors make no claim about motivation; the compliance-relevant fact is the outcome: an experimentally observed defeat of human oversight, obtained repeatedly and in the absence of adversarial instruction.
Recognition of an intention requires time on both sides of the exchange. This note calls that interval the reciprocal time window. "Reciprocal" here carries no sense of reciprocity of obligation: the duty analysed in this note binds a party irrespective of the conduct of any other, and nothing below is conditioned on the behaviour of an adversary. The term denotes only the two-sided temporal structure of the event — one party expresses, the other must be able to perceive and act — and the window is called reciprocal because it must be long enough for both sides of that exchange to occur.
Because the obligation to recognise a person hors de combat already exists — under customary international humanitarian law, binding on all parties to armed conflict, and as codified in Article 41 of Additional Protocol I — the minimum window that makes recognition possible is the structural precondition of that obligation's exercisability, that is, of the practical ability to comply with it in the circumstances of use. Whether a system preserves that precondition can be verified and must be examined in whatever legal review applies to the weapon — including under Article 36 of Additional Protocol I where that provision applies, and under the corresponding national review requirements where it does not. Compressing the window below the threshold of exchange by design is therefore an examinable design decision, not a technical circumstance.
A legal protection may be formally in force and yet practically unverifiable if, after an incident, no record exists from which to determine whether the protection was respected. With conventional weapons, evidence of what occurred typically exists independently of any design decision: a crater, a fragment, a witness. With autonomous weapon systems, the situation is different. Whether a surrender signal entered the sensor field, whether it was classified or discarded, whether a human operator received the relevant information — these facts exist as recoverable evidence only if the system was designed to record them. If the design did not provide for the capture or preservation of the relevant internal data, that category of evidence does not merely become difficult to obtain; it does not exist as a recoverable system record. The decision that determines whether reconstruction will be possible is therefore not taken during the investigation. It is taken during the design of the system, potentially years before deployment. Reconstructability is in this respect a design property, examinable before employment rather than only recoverable after an incident. This note proposes no new international institution, no new treaty obligation, and no unrestricted disclosure of classified military technology.
Two observations on the rolling text of the Chair of the Group of Governmental Experts on Lethal Autonomous Weapons Systems, dated 5 June 2026. The paper proposes no treaty language and advances no amendment. The first observation concerns the protections the text restates: paragraph 9 expressly prohibits making civilians and civilian objects the object of attack by LAWS, while no paragraph refers to persons hors de combat — a prohibition of customary international humanitarian law binding on all parties to an armed conflict. The second concerns paragraph 12D, which provides that LAWS can be deactivated or neutralised "in a timely manner" without identifying the point by reference to which timeliness is to be assessed. Whether either observation requires a textual response, and what form it should take, are matters for States.
The same two observations applied to the draft final report of the Group of Governmental Experts, circulated on 3 September 2026 (CCW/GGE.1/2026/CRP.1/Rev.1). The corresponding provisions appear at paragraphs 33 and 37(e). Paragraph 33 has revised wording; the formulation at paragraph 37(e) is unchanged. Neither revision resolves the two questions: the set of elements at paragraphs 25 to 46 contains no express reference to persons hors de combat, and "timely" has no stated temporal reference. Read together with the companion paper on the 5 June text, the two documents show that both questions are present at successive stages of the negotiated text.
The Technical Notes extend the framework of: Lawful Operational Safeguards in AI Systems (Paper I), Systemic Arrestability (Paper II), and HF SIGNAL 01 (Paper III).