Coding Agent Proof Spirals

2026-10-05 · agents

View as Markdown →

coding agentsagent designskills

A proof spiral: Build sits at the center, and each loop outward adds a larger step: check, record the check, review the record, check the checker. The spiral never reaches Ship.

I work with coding agents on long tasks where they plan and develop on their own. Most of the time it works. Every so often an agent falls into an anti-pattern I call a proof spiral. It starts writing checks of its checks.

What it looks like

The agent builds a feature and writes a check. The check needs a record. The record needs a review. The review finds a gap in the check, so the agent builds a better checker. Each step looks reasonable. Together they stall delivery, and the logs look rigorous the whole time.

These are the signs I watch for:

Watch the status lines too. When they fill with words like frozen, sealed, byte for byte and exact receipt, and keep ending in “remains pending”, I go look.

Why it happens

It’s tempting to blame the model. Look at the workflow first.

Workflow rules often make defects and violations explicit. They rarely say how long delivery should take or how much uncertainty is fine to leave behind. Take three reasonable rules: “always verify independently”, “fix every finding” and “persist until done”. Together they prevent completion. A model that follows instructions well obeys all three. Strong instruction following becomes the liability.

Rules also pile up. Each bad run earns a new rule, and the old rule stays.

Checking can also stand in for access. An agent that can’t reach the real environment builds mocks and checkers instead. More local checks can’t settle an external question.

How to stop it

Start with one question: What decision does this proof change? If the answer is none, the proof is ceremony. In the spirals I’ve seen, this question would have stopped them earliest.

Then fix the workflow:

Heavy checking isn’t always a spiral. Releases, migrations and security changes can need it. Keep the checks that earn their place: behavior tests, one review, the release rehearsal, and checks on money, data, credentials and users.

The skill

I wrote this up as a skill you can drop into your own agent setup: proof-spiral.md. It covers how to notice, diagnose, address and prevent a spiral. It also warns about itself. The skill must not become one more layer of rules.

Also: introducing ThinkThen

I recently introduced ThinkThen. It answers typed questions about text and returns true, false, a label or a number. A failed call never looks like an answer. Use it to gate a script, label records, or grade answers in an eval.

thinkthen decide 'Does the customer ask for money back?' < message.txt

The command prints true and exits 0 for yes, 1 for no and 3 for not sure. A shell if can branch on it. It has ten functions, such as decide, choose, tag and score. They run across 25 surfaces: the command line, a Rust crate, language bindings from Python to COBOL, and SQL extensions for DuckDB, SQLite and PostgreSQL.

Watch on YouTube →

Read the docs at thinkthen.dev. The code lives at github.com/botassembly/thinkthen.