Jeremy Nelson / Chicago / local AI / agent systems

Make the system legible.

I build and test practical AI agent systems on real machines. The work is routing, state, evals, local inference, failure visibility, and the evidence trail that makes an autonomous run inspectable after the impressive part is over.

agent-run.log live

$ objective: make uncertainty smaller

$ boundary: scoped tools, visible state

$ proof: files touched, checks run, failures named

$ local: fast loops beat sovereignty theater

$ stop: when the next safe step is obvious

position AI systems engineer
surface agents under pressure
taste proof over performance

Evidence beats narration.

Useful agents leave a run record: goal, repo state, commands, failures, checks, diff, and the reason the work stopped.

Local AI is a loop-speed tool.

The value is a shorter path from patch to test to recovery when the machine is sitting on your desk.

Harnesses should preserve weirdness.

A good harness captures the edge case, replays the failure, and makes the next model prove it did not just get lucky.

Memory is not operating state.

The agent can be disposable. The work record cannot. Context needs source, freshness, scope, and a reason to go quiet.

Routing is product judgment.

Model choice should expose task type, tool access, fallback reason, cost, latency, and the point where a human can override it.

Failure visibility is the UI.

The useful workbench shows current objective, files touched, failed action, evidence changed, and the next safe step.

My lane is the unglamorous layer between a demo and a system people can trust: permissions, state, replay, tool boundaries, eval traces, model routing, local hardware constraints, and the human-readable artifact left behind.

I write from the machinery. The best agent work feels like a clear workbench: current state, visible edges, failed attempts, and proof that the next move is safe.

Things I keep coming back to

  • cost per verified outcome
  • run ledgers and handoff quality
  • local model latency and recovery loops
  • tool permissions, blast radius, and rollback proof
  • the carrying cost of generated code