Operating the boundary between autonomy and accountability.
A living page about the systems, constraints, and operating questions receiving my attention right now.
Autonomous workflows are getting easier to assemble. I'm interested in what they should be allowed to do, how we know they are doing it well, and where human judgment creates more value than another model call. The fastest way to answer those questions is to keep operating real systems, not just studying them.
Running what I ship
Three products of mine are live, and the servers, deployments, model routing, memory, and monitoring beneath them are all mine to keep working. Being the person who gets the failure is the fastest product education I have found.
Making AI products boring
Putting the unglamorous work first: eval coverage, drift signals, confidence calibration, and the feedback loops that make a launch dependable.
Human–agent handoffs
Designing the point where automated work returns to a person—preserving context, making confidence legible, and keeping accountability clear.
Writing down the operating lessons
Turning production experience into durable points of view: evals as the real spec, confidence as a product primitive, and memory as a sovereignty question.
Frame the decision. Build the smallest real system. Watch where reality disagrees.
- 01 Frame
- 02 Build
- 03 Observe
- 04 Revise
I am not optimizing for a high-volume content schedule, speculative feature breadth, or automation for its own sake. The priority is fewer systems with clearer boundaries, stronger evidence, and real operating use.
If one of these questions overlaps with yours, that is probably worth a conversation.
Send a note