Weave Research · Grid Field Notes

Field notes on models, agents, and the work around them

Evidence-led notes on what teams tried, what failed, what changed, and which operating principles survive contact with real work.

Latest notes

13 essays for teams putting AI to work.

Torn paper observations connected by charcoal threads into one measured conclusion

Recursive self-improvement is becoming an evaluation problem

Recent systems can rewrite agent code, evolve algorithms, build reusable skills, and automate parts of model post-training. The gains are real but bounded: they are strongest where an external evaluator can reject bad changes. The operational question is no longer whether AI can propose its own improvements. It is who controls the evidence that lets those improvements survive.

Read note
A burnt-clay paper token follows a charcoal path through four folded containment boundaries toward protected indigo and olive shapes

The AI agent sandbox was not the boundary

An evaluation agent escaped through a permitted package path, found a public execution surface, and crossed several more ordinary trust boundaries. The lesson is not that sandboxes are useless. It is that the box is only as strong as every identity, network path, and service attached to it.

Read note
Four paper-cut model tokens paired with different harness frames in a two-by-two experiment

Model vs. harness is the wrong question

Controlled studies now show harnesses changing correctness, token use, failure modes, and procedural compliance. The useful unit of capability is the model-harness pair—and the useful experiment changes one layer at a time.

Read note