professional headshot photo of Mr Palumbo

The Forward Deployed Audit: Deciding Which Workflows Actually Need AI

Forward deployed engineering is usually pitched as installation work: an engineer shows up, wires a model into your systems, and leaves you "AI-powered." That pitch has the job backwards. The real deliverable is a classification. You walk the business's actual workflows step by step, and for each step you answer one question: does this need judgment, or does it need a script? Most steps need the script.

I run software this way for a living - for a paying client, and for my own operation, a portfolio of production apps maintained largely by autonomous agents. Both were audited with the same method, and both landed on the same shape: deterministic software does the bulk of the work, agents are reserved for the narrow set of steps that need a reader, and a human holds the irreversible ones. Every number below is measured from systems I operate.

204
deterministic tools
56
judgment skills
2 of 4
scheduled jobs need a model
~96%
of tokens are cache reads
01 · The job

Deployment is an audit, not an installation

A forward deployed engineer embeds with a business in a way a normal vendor never does: you sit inside the workflows, run them end to end with the people who own them, and then own what gets built. The tempting move - and the sales pitch of the moment - is to route everything through a model, because a model can plausibly do anything. Plausibly is the problem.

Non-determinism has a price list. Tokens cost money on every run. Latency is variable. The same input can produce different outputs, so every output needs review, because it can be differently wrong each time. Paying those prices for work a script could do identically forever is the defining mistake of enterprise AI adoption right now. The inverse mistake is quieter but just as expensive: scripting a step that needed a reader, and silently mishandling every case the rules never anticipated. The audit exists to keep you from making either one.

02 · The audit

One question per step

The method is unglamorous. Take one workflow at a time and name its output first - a workflow whose output nobody can name is a wish, not a workflow, and no automation of any kind will fix that. Then decompose it into steps and split each step's deterministic half from its judgment half.

The operating rule I hold my own automation to is written into my organization's manual: "Deterministic scripts gather; agents spend tokens only on judgment over script output." The classification rule follows from it: if it acts with judgment on a trigger, it is agent work; if a re-run must produce the identical result, it is a tool. A tool that makes judgment calls cannot be trusted to re-run. An agent that re-derives arithmetic burns money doing it.

A worked example from my own systems: triaging user feedback. Mapping a submission tagged bug to a high-priority work item is a rule - it must not drift between runs, so it is a script. Deciding that "the pricing page confused me," typed into a plain feedback box, is actually a real work item worth scheduling - that needs a reader. The script files the obvious cases and holds the ambiguous ones; the agent reads only what was held, and must state a reason for every promotion. Same workflow, two halves, priced separately.

03 · The framework

Scripts, agents, humans

Classifying a step takes three questions, asked in order. Irreversibility is checked first, because it overrides everything else.

The audit decision flow One step of a business workflow Does it touch auth, payments, a schema, secrets, or user-data deletion? yes Human gate always - no stage skips it no Must a re-run produce the identical result? yes A script cheap, testable, identical no Does the step need a reader? judgment over facts a script gathered yes An agent tokens only for judgment no Leave it with a human batch it, and revisit when it recurs
The audit, one step at a time. In practice most steps exit at the second question - that is the whole thesis in one picture.

Inside my own systems, the deterministic column is deliberately large. These stay scripts by design, even though an agent could do them:

  • Gathering. Surveys of project state, exports, and syncs run as scripts that print structured output. The standing instruction is literal: run the script, do not reason.
  • Verification. Tests, linters, type checks, and security scans decide pass or fail with no model in the loop. A model reviewing its own work is not verification.
  • Routing. Which class of model handles a task is a keyword match on the task's text, not a model call - the same cheap determinism, pointed at cost and scheduling.
  • Deduplication. Re-runs are made safe by idempotency keys, not by a model comparing items and guessing what it has seen before.
  • Status. Verdicts like failing, stale, or never-run come from hard thresholds, so they are identical every run - and a dead data source degrades to unknown, never to ok.

The judgment column is short, and every entry on it involves reading: reading a customer's free-text feedback, writing the specification for a change (with acceptance criteria a machine can check afterward), verifying that a task's premise is still true before building it, and naming a failure mode nobody has seen before. That is what the tokens are for.

Then there is the third column. Some steps are neither script work nor agent work at any level of maturity:

Agents doHumans hold
Create branches and isolated workspacesDefine intent and acceptance criteria
Write code, tests, and documentationApprove or reject the change
Run builds, tests, linters, and scannersMerge to protected branches
Open and update pull requestsGrant, scope, and revoke credentials
File the follow-up work they discoverChange the guardrails themselves

And a set of hard carve-outs overrides everything above: any change touching authentication, payments, a schema migration, secrets, or deletion of user data routes to a human first - always, no matter how mature the automation around it has become. Blast radius has no stage.

04 · The case study

A shop with thirteen seasonal staff

The client is a seasonal deer processing and taxidermy operation in rural Indiana: a purpose-built facility, about thirteen seasonal staff, and a few frantic months a year when hunters line up at the counter. The engagement was to take check-in digital - the intake forms for processing and taxidermy orders, and everything the shop does with them afterward.

This is exactly the kind of business the current wave wants to sell a chatbot to. The audit said otherwise. Walking the workflow with the owners produced a system with no model in it at all:

The client check-in system Online check-in forms processing and taxidermy Database every record starts pending Staff pending queue the inbox - no emails sent staff work the queue at the counter Official check-in animal physically present Money fields + notes staff-editable only Audit trail who changed what, when CSV export opens in Excel Per-submission owner emails: none. The pending queue is the inbox. Email is the transport, never the tracker. No model in the request path.
The delivered system. Notice what is missing: there is no model box, because the audit found no step that needed one.

The interesting decisions are the negative ones, and each came from listening rather than from a feature list. An online submission is not an order - it becomes official only when staff check it in at the counter with the animal physically present, because possession is the shop's real source of truth. The owners get no notification email per submission, because during the season that would be a hundred emails a day nobody reads; the pending queue is the inbox, worked at the counter. Customer order details are deliberately not editable in place: small changes go on butcher notes, big changes go through a fresh check-in cross-referenced to the old one, and the money fields staff actually need to adjust are the only fields staff can adjust - with the audit trail recording who touched what. Exports are CSV because the back office runs on Excel, not on dashboards.

Even the client communication followed the audit's logic. Questions went out in written batches - seven rounds so far - published as web pages first, because a page can be translated by any phone in one tap, with a PDF fallback that has ruled answer lines, because shops print things and write on them. Structured-enough answers, zero new software for the client to learn.

Here is the punchline. The AI content of the delivered product is nothing: no model in the customer-facing request path, no chatbot, no "AI-powered" anything. Every flow the audit surfaced needed deterministic software that behaves identically on the ten-thousandth deer as on the first. The AI leverage in this engagement was in the delivery - agent-driven builds working from written specifications, behind the same review gates described above. The client bought outcomes, and the audit is what kept them from buying tokens instead.

05 · The evidence

The ratio holds at fleet scale

One client engagement is an anecdote. So I ran the same audit on my own operation - a portfolio of 69 active repositories maintained by one human plus autonomous agents. Over a 41-day window the system ran 103 unattended passes, produced 1,001 work-item rows across 75 projects, and verifiably shipped 342 of them (a floor - the real number is higher, but I only count what leaves a trace). July alone saw 2,089 commits against a pre-automation baseline of roughly 40 a month. This is not a demo; it is the most agent-saturated environment I can measure. And even here, the automation surface is mostly scripts:

Deterministic tools vs judgment skills Shell tools 166 Scripts 38 Judgment skills 56 204 deterministic (166 shell tools + 38 scripts) to 56 judgment skills - about 3.6 : 1
The automation surface of my own organization, counted from the repository. The highlighted bar is the minority: the parts that spend tokens.

The clearest artifact of the audit is the scheduling manifest, where every recurring job carries a machine-readable answer to the question "does this need a model?" Half do not, and that answer has an infrastructure consequence: the deterministic jobs can migrate to any commodity Linux host, while the agentic ones are pinned to a machine with model access.

Scheduled jobKindNeeds a model
Weekly repository auditdeterministic scriptno
Monthly bookkeeping roll-updeterministic scriptno
Daily roadmap passagenticyes
Weekly portfolio digestagenticyes

Even where the agents do run, the spend shape is engineered like infrastructure rather than magic: of 11.82 billion tokens metered over 29 active days, about 96% were cache reads - repeated context served from cache at a fraction of the cost - with only 27.9 million tokens of actual model output. The same platform also ships models where they earn their place: consumer products of mine like Lost Pet Radar and WriteMyCard AI put image moderation and LLM drafting in front of paying users, behind the same deterministic billing, auth, and deploy rails. The audit is not anti-AI. It is anti-waste.

06 · The takeaway

What to buy, what to build

If you are a business considering AI, the audit reframes the purchase:

  • Buy judgment only where a step needs a reader. Free-text, ambiguity, novelty, and specification are agent work. Everything with a nameable, repeatable output is a script wearing an AI costume.
  • Prefer deterministic software for the bulk. It is cheaper per run by orders of magnitude, it is testable, it is auditable, and it does not degrade or drift while you sleep.
  • Keep the irreversible steps human. Auth, money, schemas, secrets, and deletion do not get automated judgment at any maturity level - the cost of being wrong once exceeds the cost of every approval you will ever click.
The takeaway The deliverable of a forward deployed engagement is not a model integration. It is the classification: this half of your business is scripts, this sliver needs judgment, and these steps stay yours. Most of what the audit classifies is a script - and an engineer who tells you that is worth more than one who sells you tokens.