Notes from production
Field notes on getting AI agents past the demo and into real operational work, from someone who builds and runs them.
Pull Three Timestamps Before You Buy Any AI Tool
Most AI purchases fix the wrong stage. Three timestamps per ticket from last quarter will show you exactly where work actually waits.
Run a Manual Work Inventory in One Week and Know What to Automate First
Most automation roadmaps fail before they start. A structured inventory built in five days fixes that by surfacing the real cost of repetitive work.
You Automated the Bottleneck and the Constraint Moved
The first agent ships and throughput jumps. Then the next stage chokes. Here is why that is predictable and how to get ahead of it.
Measure Queue Time Not Touch Time to Find Where Work Actually Stalls
The minutes a person spends on a task are rarely the problem. The hours it sits waiting are. Here is how to instrument the waiting.
Process Mining vs the Two Hour Workshop Which One Finds the Real Bottleneck
Event logs and team interviews answer different questions. Knowing which to run first saves months of analysis and avoids the most common implementation mistake.
Your Developers Are Not the Bottleneck in Software Delivery
Review queues and environment waits consume more calendar time than coding in most pipelines. Here is how to see that in data you already have.
Five Signals Your Workflow Has Friction Before Anyone Complains
Complaints are lagging indicators. These five patterns show up in your operations weeks before anyone files a ticket or raises a hand.
Four Ways to Find Bottlenecks in Operations Workflows and What Each One Actually Costs You
Process mining, ticket analytics, time studies, and queue-reading agents each surface different problems. Here is what they find, what they miss, and what it takes to stand them up.
Explainability in Production AI Agents Is an Operational Requirement Not a Feature
What a cited, reviewable recommendation actually looks like in a working system, and why the extra engineering pays back fast.
AI Governance That Ships Fast Without Gutting Your Controls
Most governance frameworks stall AI programs before they deliver value. Here is how to build controls that satisfy risk and still let your team move.
Map the Manual Work Before You Build the Agent
The teams that ship working AI agents start with a spreadsheet audit, not a prompt. Here is the order that actually works in production.
Edge Cases Are Where Production Agents Live or Die
The unusual five percent of inputs your agent was never trained on will determine whether operations trusts it or abandons it. Here is how to find them first.
What an AI Agent Actually Costs to Run in Production
Token spend is maybe 15 percent of your real bill. The rest is evaluation, monitoring, integration, and the human work you forgot to budget for.
Your First Ninety Days With AI in Operations A Sequence That Actually Works
Skip the pilot theater. Here is the order of moves that turns one boring win into a platform your team will actually use.
Trust Beats Accuracy When Deploying AI Agents in Operations
A system your team believes in will outperform a smarter one they route around. Here is how to earn that belief in production.
When a Script Beats an Agent and You Should Say So
The most credible thing an AI engineer can tell a client is when not to use AI. Here are the cases where simpler wins every time.
The Regression Test That Caught a Payroll Error Before It Shipped
A replay suite over real past runs will catch what a unit test never sees. Here is what that looks like when the stakes are actual paychecks.
What a Production Agent Owes the Team That Runs It
Shipping an AI agent is easy. Keeping one running safely under real operational pressure is a different contract entirely.
The Prompt Is Not the Product Why Production Agents Live Elsewhere
Everyone obsesses over prompt wording. The engineers who actually ship agents know the prompt is maybe five percent of the work.
Build Buy or Wrap How to Choose the Right AI Approach for an Operational Workflow
Most teams pick the wrong option not because they lack options but because they ask the wrong question first.
What Actually Breaks When Your AI Agent Scales from 10 Runs to 10000
The failure modes that kill production agents at scale are almost never the ones you tested for. Here is what changes and what to do about it.
Why a 95 Percent Accurate AI Quote Will Destroy Customer Trust
One wrong answer in twenty is not a rounding error. In customer-facing AI, it is a reputation event that compounds faster than any speed gain can offset.
Five Questions Every Finance Team Should Ask Before Trusting an AI Recommendation
The difference between a decision-support agent and a black box is not the model. It is whether your team knows how to interrogate the output.
Automate Field Reporting and Keep Every Number Traceable to Source
Parallel data gathering and structured reconciliation let AI agents produce reports fast without breaking the audit trail that regulators and ops leaders depend on.
Scope Your First AI Agent to One Workflow and Ship It Fast
Most enterprise AI projects stall because the scope is too wide. Cut to one workflow, one decision, one outcome and you will ship in weeks.
One Agent or Many Agents How to Decide Before You Build
Multi-agent systems earn their complexity in specific conditions. Most teams reach for them too early and pay for it in debugging time and drift.
Your First AI Win Should Be Boring and Repetitive
The highest-value first AI project is never the impressive one. It is the dull, high-volume task your team does the same way a thousand times a week.
Traceability Beats Fluency Why Sourced AI Gets Adopted in Regulated Work
A confident answer nobody can verify kills adoption faster than a wrong answer. Here is how grounding every claim to a source changes that.
Automating Order Entry Is a Data Problem Before It Is an AI Problem
Catalogue mapping and line matching decide whether your order automation works or fails. Here is how to scope it so accuracy is provable before you go live.
How to Measure Whether an AI Agent Actually Saved Time
Most teams celebrate automation before they measure it. Here is how to baseline honestly, count correctly, and report a number finance will believe.
Put a Rules Engine Next to Your LLM Before You Regret Not Doing It
Deterministic logic and language model judgment are not rivals. Knowing which one owns each decision is the whole game in production AI.
The Model Is Not Why Your AI Project Stalled
In production AI work, the model is almost never the bottleneck. Here is what actually kills projects and what to fix first.
What an AI System Owes You That an AI Feature Never Will
Shipping a model into production is not the same as running a system. Here is what separates the two, and why it matters when things go wrong.
Reading a Document With an Agent Is Easy Trusting It Is the Hard Part
Extraction accuracy is table stakes. The checks that turn a document reader into something operations will actually sign off on are a different problem entirely.
Safe Failure Is a Design Choice Not an Accident in Production AI Agents
The plumbing that separates a bad agent run that fixes itself from one that corrupts three days of inventory records.
Human in the Loop Is a Design Decision Not a Disclaimer
Where you put a human in an AI workflow determines whether the system learns and earns trust, or just shifts liability around.
Evals Are Why Anyone Trusts Your Agent in Production
Replay regression, not gut feel, is what separates an agent a team will actually use from one they quietly route around.
Why AI Pilots Look Great and Then Die Before Production
Most AI pilots fail not because the model is wrong but because nobody built the system around it. Here is what actually separates demos from deployable agents.