SHEET 01 · WRITING
Essays on AI, agents, and harness engineering.
Field notes from building enterprise GenAI at production scale. The rise of harness engineering, why prompts are technical debt, how AI coding agents are reshaping software work, and the convergent moves of the major labs. Read each piece here on the site, or on Medium.
- AI Agents2026-10-06 · 12 min read
What Claude Code Keeps of Your Skills After Compaction
A 5,000-token snapshot, a 25,000-token shared budget, and a skill index that never comes back. The mechanics, the failure modes, and the patterns that hold up in long sessions.
- LLM Systems2026-09-27 · 9 min read
Building Complex AI Workflows with Jev
A live experiment in composing small judgments into larger decisions and knowing where the approach stops being useful
- LLM Systems2026-08-25 · 20 min read
MCP 2.0? What Changed, What Remained, and How it Impacts Practitioners
Working notes on the 2026-07-28 revision of the Model Context Protocol, from someone responsible for running it in a large, regulated enterprise. Written for peers carrying the same responsibility.
- AI Agents2026-07-27 · 17 min read
Your Agent Has Too Many Tools, and No Way to Take Any Back
Claude Opus 5 shipped with two beta features listed at the bottom of the announcement. One of them changes where an agent’s authority is allowed to live.
- LLM Systems2026-07-26 · 9 min read
Open Weights Are Good Enough. The Hard Part Is Everything After.
They now match commercial models on most enterprise work at a fraction of the cost. The differentiating skill is no longer picking a model. It is drawing the open-versus-commercial line well, and running the open side with discipline.
- LLM Systems2026-07-12 · 12 min read
GPT-5.6: What Actually Changed on Your Bill
Read the diff, not the demo. Field notes on the GPT-5.5 to 5.6 migration, where the flagship price held but a few things that used to be free now cost money and several defaults moved underfoot.
The Agentic Engineer
- Part 1
The Agentic Engineer [Part 1/3]: Claude Code Is Not a Chatbot. It's an Engineer You Build Structure Around.
LLM Systems2026-06-21 · 7 min readWhy the structure is the job, the prompting is secondary, and how I run an AI coding agent like a production engineering practice.
- Part 2
The Agentic Engineer [Part 2/3]: Shipping a Feature Without Losing the Thread
LLM Systems2026-06-21 · 8 min readThe full pipeline I run when a task is too big to hold in my head: spec, data contracts, controlled fan-out, arbitrated review, and bounded autonomy. This is the heavy lane, and it only earns its cost on real features.
- Part 3
The Agentic Engineer [Part 3/3]: Into the Brownfield
LLM Systems2026-06-22 · 7 min readParts 1 and 2 assumed greenfield, where I set the conventions. Most real work isn't that. This is what changes when you point an AI agent at a large codebase you didn't write, where the rules already exist and the blast radius is unknown.
- Part 1
- AI Agents2026-06-12 · 10 min read
Decompose First, Judge Last
Field notes on evaluating LLM systems in production: what actually catches failures, what quietly doesn't, and why most teams reach for the wrong tool first.
- AI Agents2026-06-08 · 6 min read
Everything Your AI Agent Reads Is Executable
Inside an AI agent, the line between data and instructions disappears. That single blur is the root of nearly every security risk in agentic AI.
- AI Agents2026-06-06 · 4 min read
The LLM API Call Quietly Became an Agent Loop
One request now runs a server-side loop of model passes and tool calls. Here is what that buys you, and what it quietly costs.
- AI Agents2026-04-18 · 16 min read
Anthropic and OpenAI Just Shipped the Same Answer to AI Agents, Seven Days Apart
Inside a seven-day window in April 2026, the two largest frontier labs independently shipped the same agent architecture. The convergence matters more than either release.
- LLM Systems2026-04-13 · 11 min read
Your Prompts Are Technical Debt: A Migration Framework for Production LLM Systems
Your prompts aren't instructions. They're contracts with an expiration date.
- AI Agents2026-04-10 · 10 min read
Your AI Coding Agent Is Winging It. Superpowers Makes It Stop.
The most popular Claude Code plugin doesn't make your agent faster. It makes your agent think before it types.
- AI Agents2026-04-08 · 8 min read
Making the Most of Claude Code: Best Practices & Hidden Features
A practical guide to the underused features, productivity workflows, and best practices for getting maximum value from Claude Code.
- AI Agents2026-04-07 · 7 min read
The Rise of Harness Engineering: Why the Code Around the Model Matters More Than the Model Itself
One team moved a coding agent from the bottom 30 to the top 5 on a leaderboard. Same model, same weights, zero retraining. They only changed the harness.
SHEET 02 · GET NEW ESSAYS BY EMAIL
Get new essays by email
Get an email from Medium whenever I publish.