Skip to content

Research & Projects

Most of what I learn about agents comes from running them in production, and when I can, I turn that into something other people can check and reuse. Everything below is public, with the methodology written down, including the parts that did not work.

Open research

LHC: a benchmark for long-horizon agent coherence (2026)

An open benchmark for whether 8B-class models keep track of state, commitments and half-finished work across long stretches of unrelated context. It has 24 hand-curated tasks across three failure modes, each run under four gap conditions, and a decision matrix that was locked before any model was scored. The release also includes a deterministic parser of about eighty lines of Python, which turned out to be a useful floor for structured-state tasks, and the full journal of a fine-tune that did not measurably beat its base model. I published that result rather than the model.

Dataset on Hugging Face · Code · Parser baseline · Write-up

How agents fail in production: MAST on real telemetry (2026)

Most failure research studies benchmark tasks. I applied the MAST taxonomy to 639,381 execution steps across 23,624 runs on a closed-alpha agent platform over five months. The most useful result was methodological: my first headline finding was an infrastructure bug that looked like agent behaviour. The code and aggregate results are public, and every number in the write-up can be reproduced from them.

Code and results · Write-up

Production systems I write about

At Complyance I built the agent runtime our compliance agents run on, and a broker that lets them call customers’ GitHub, AWS and other systems from a sandbox that never holds a credential. That code is not open source, so I write about the architecture in as much detail as I can:

Companies and ventures

More code lives on GitHub.