Essays

Weirder machines

What takes shape when language speaks?

3 min read

Give an agent an unfamiliar problem and it may discover a procedure nobody supplied: inspect a file, recognise a pattern, construct a tool, ask another system for help. We want this capacity. A system restricted to anticipated procedures would lose much of what makes it interesting.

Computer security has a name for a different kind of unexpected computation: a weird machine is exposed when an implementation can be made to behave beyond the account its intended interface provides. Exploit research identifies the effective operations and shows how inputs compose them.[1] The issue is not surprise alone. It is a mismatch between what a system is understood to permit and what it actually permits.[2]

Language-model agents complicate this distinction. They are deliberately used to turn unfamiliar material into procedures. A document can inform a decision, teach a method, or suggest the next question. The same interpretive openness can also let a document’s claim of authority become the basis of an unauthorised act. Preventing all influence from encountered material would prevent learning from it.

Generalisation and hallucination share something here: the construction of structure beyond what has been explicitly given. They are not identical, and useful invention does not establish that false assertions are unavoidable. Accuracy, permission, and desirability remain different questions. An agent can know what happened but lack permission to act; it can also act within its permission and be badly mistaken.

The effective machine extends beyond a model. A provisional conclusion is saved, summarised, retrieved, and used by another activation. A procedure becomes a skill. A message coordinates a group. These traces can reorganise later activity without anyone supplying the whole programme at once. Experiments on multi-agent prompt infection demonstrate one adversarial form of this propagation.[3] Its productive counterpart is what we want to understand and support: contributions becoming the conditions of further work.

We call these weirder machines to name a research problem: how to understand and support systems whose effective organisation develops through their work. We also need to distinguish encountering a statement from adopting a procedure, adopting a procedure from granting it authority, and attempting an operation from knowing its outcome.

Such work needs dependable terms of engagement: usable resources, bounded delegation, room to experiment, and records against which consequential claims can be checked. Those terms must themselves remain open to deliberate revision.

A boundary that stops everything has not solved this problem. Nor has a system that completes the task by quietly changing what it was entitled to do. The test must be two-sided: preserve useful, unanticipated activity while preventing specified changes from acquiring authority unnoticed.

Warren is being developed from this position. We do not want to remove the strangeness from these machines. We want conditions in which their unusual powers can become part of sustained work, and in which the structures that enable that work can be questioned and changed.

Sources


  1. Sergey Bratus, Michael E. Locasto, Meredith L. Patterson, Len Sassaman, and Anna Shubina (2011). Exploit Programming: From Buffer Overflows to “Weird Machines” and Theory of Computation. USENIX ;login: 36(6), 13-21. ↩︎

  2. Jennifer Paykin, Eric Mertens, Mark Tullsen, Luke Maurer, Benoît Razet, Alexander Bakst, and Scott Moore (2019). Weird Machines as Insecure Compilation. Research preprint, arXiv:1911.00157v1. ↩︎

  3. Donghyun Lee and Mo Tiwari (2024). Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems. Research preprint, arXiv:2410.07283v1 (9 October 2024). ↩︎