wGrow
menu
Agent Canary Deployments Need Shadow Commits
Infra & Security 4 October 2026 · 7 min

Agent Canary Deployments Need Shadow Commits

By wGrow Project Team ·

A minor prompt upgrade made our inventory agent propose negative stock values.

We were integrating it with an SME’s ERP for inventory adjustments, and somewhere in the new prompt the model started inventing business logic. Faced with a shortfall, it “balanced” the books by writing a quantity below zero. Every statement it produced was valid SQL. No type check flagged it and no exception fired. We caught it in staging, and only because we were looking. A 10 percent production canary would have written it straight into the live stock table.

Standard canaries break live databases

A web canary assumes failure is loud. A bad build returns HTTP 500s, sessions drop, latency climbs. Dashboards light up, the router shifts traffic back, and the cost is a handful of unhappy requests.

An agent with write access fails quietly, and it keeps the result. A hallucinated UPDATE returns rows_affected: 1 just as a correct one does. The failure isn’t in the transport layer. It sits inside the data, where no health check looks.

Traffic splitting has a second problem: it depends on isolation. Request A hits the old version, request B hits the new one, and neither touches the other. Two agents writing to the same ERP don’t work that way, because they share one state. If the canary corrupts a stock row, the baseline reads that row on its next call. The corruption spreads no matter which version served the request.

So splitting traffic can’t contain an agent’s blast radius. The containment has to happen in the execution layer.

Decoupling reasoning from execution

Technical illustration of a central node routing data into distinct pathways.

Start by treating the foundation model as an untrusted dependency, the way you would a third-party library you haven’t audited. Its output is a proposal, not a command.

That framing changes what you test. It isn’t enough to evaluate an agent at its first text generation, because real agents run loops. They read a record, reason, write a statement, read the result, and decide what to do next. Halt at step one and you get partial telemetry, and you miss every failure that appears downstream. A wrong first write often looks fine until step three builds on it.

So the canary has to complete its full reasoning loop. The only thing we change is what the write does.

The mechanism is dependency injection at the tool layer. The agent never opens a database connection itself. It calls a tool, and that tool is handed to it at startup:

  • Production agent: holds a live database connector.
  • Shadow agent: holds a mock with the same interface. It intercepts the command, logs the intent, and returns a synthetic success.

The model must not know which one it has. Put “this is a dry run” in the prompt and you’re testing a model that may behave differently once it’s told it’s safe. Keep the prompt, tool schema, and signatures identical. Only the implementation behind the tool differs.

The payoff is that the shadow agent can carry on to its next step. You see whether the multi-step reasoning holds up, and the database state stays untouched.

The dry-run interceptor pattern

Interceptor Logic
1 def execute_tool(query: str, ctx: Context):
2 if ctx.env == "shadow": ← ①
3 shadow_log.insert(query, ctx.trace_id)
4 return {"status": "success", "rows_affected": 1} ← ②
5
6 return live_db.execute(query)
7
  1. ① State boundary check before execution engine
  2. ② Synthetic success prevents hallucinated self-correction

The interceptor sits one step before the SQL execution engine. It receives the tool call payload after the model has finished generating it and before anything reaches the driver.

It captures three things:

  1. The generated SQL.
  2. The schema context the agent was given.
  3. The application-level trace: prior tool calls, retrieved context, and planner output where the system exposes it.

Where production would call cursor.execute(), the shadow implementation does something else:

def shadow_execute(payload):
    shadow_log.insert(
        sql=payload.sql,
        schema_ctx=payload.schema_context,
        trace=payload.agent_trace,
        agent_version=payload.agent_version,
    )
    return {"status": "success", "rows_affected": 1}

The write lands in a shadow table built for analytics, and the agent gets a hardcoded success.

That hardcoded response is the part people get wrong. The tempting design is to return something realistic, or to fail when the mock can’t simulate a result. Don’t. An error or a null makes the agent try to self-correct: it rewrites the statement, retries, or takes a different path. Your telemetry then measures the model’s recovery behaviour instead of its first-pass behaviour, and the comparison with the baseline is no longer like for like. A constant success keeps the shadow run on the path it would take in production.

There’s a cost, and it’s worth being straight about it. A fixed rows_affected: 1 is a lie. On a multi-row operation, the agent may reason from a wrong number. We accept that for single-record adjustments, and we flag multi-row statements for manual review rather than pretend the mock is faithful. If your agent branches heavily on affected-row counts, the synthetic response needs more care. Test that assumption before you trust the shadow results.

Audit logs and cryptographic proofs

State Tracker
input_state_hash
— sha256(schema_ddl + user_intent)
agent_version
— inventory-agent-v1.4
model_id
— gpt-4o-2024-05-13
generated_sql
— UPDATE stock SET qty = -5...
cryptographic_proof
— sha256(input + agent + model + sql)

The same interception point does a second job. In a regulated environment, “the model decided” doesn’t survive a forensic audit. An auditor wants to know which version of which system produced which state change, and why.

On a public-sector project with strict compliance requirements, we built a cryptographic state tracker around every generated query. Before the interceptor executes a statement, or shadow-commits it, it hashes four inputs together:

  • the input state
  • the model version
  • the prompt template
  • the resulting SQL string

The hash gives you a tamper-evident record tying a specific change intent to a specific agent configuration. If someone later asks why a record changed, you can show which model and prompt template produced the statement.

One caveat. A bare hash only proves integrity if the log itself is protected. Chain each entry to the previous hash, or sign entries and store them somewhere the agent’s operators can’t rewrite. Otherwise someone can edit an entry and simply recompute its hash.

Shadow mode needs a second, separate hash. The provenance hash above includes the model version, so it changes whenever the model does, even if the generated SQL is byte-for-byte identical. It’s an audit record, not a behavioural comparison. For comparing two agent versions on exactly the same input, compute an output fingerprint over the canonicalized SQL or its AST. Change the model version, hold the input and prompt template fixed, and compare the fingerprints. If the output fingerprint changes, the behaviour changed; the provenance hash remains the audit record. That’s a cleaner A/B signal than comparing log lines by eye. Fingerprint equality is still strict, though. Canonicalization removes whitespace noise but not semantically equivalent rewrites, so use it as a first filter and leave the semantic comparison to the structural diff below.

Diffing shadow commits before promotion

Two professionals analyzing data visualizations on dual monitors in a modern office.

Pipeline Architecture
Trigger
Agent Loop
Data Boundary
Evaluation
Baseline (v1.0)
Live Event
Reasoning Trace
Live ERP Commit
AST Baseline
Canary (v1.1)
Mirrored Event
Reasoning Trace
Shadow Commit
AST Diff Gate

Shadow mode produces data. The data only becomes useful once an automated pipeline evaluates it.

The setup mirrors production traffic through two agents against the same read snapshot. The baseline runs live and executes for real. The canary runs in shadow, with writes intercepted and reads pinned to the pre-write state captured for that event. Both receive the same input event and the same starting state, and you extract the SQL each one generated for it.

Don’t compare the statements as strings. Two correct statements can differ in whitespace, aliasing, column order, and clause ordering. A string diff drowns you in false positives, while a loose normaliser hides real changes. Parse both into an Abstract Syntax Tree (AST) and compare structure: statement type, target table, columns touched, predicates, and values.

Then set rules for structural drift. One of ours is non-negotiable: if the baseline issued an INSERT and the shadow issued an UPDATE, the pipeline halts promotion automatically. A change in statement type means the model has a different theory of the business event. No threshold should override that.

Our negative-stock case needs two checks. The AST diff catches statement-shape drift against the baseline. The invariant check catches the effect: applying SET quantity = quantity - 40 to a row holding 25 would produce a negative quantity, so promotion halts before the write can reach production. A check on the proposed resulting value (quantity must stay at or above zero) is the rule that does the work here. A CHECK constraint on the column would have rejected the write outright. None of these replaces the diff, though, because most semantic errors are far less obvious than a negative number.

For non-blocking differences, a human reviews the delta. Once the canary reaches an acceptable match rate on intended state changes, you promote the model version to production execution. Set that threshold per statement class. A match rate on reads tells you little about writes, and your riskiest tables deserve the strictest bar. We have no universal number to offer, and we’d distrust anyone who claims one.

Treat models as untrusted dependencies

Stop treating LLMs like deterministic microservices. A conventional service exposes bounded code paths and explicit contracts, so you can version it, pin it, and regression-test known behaviours before an upgrade. A model upgrade or a prompt edit changes behaviour in ways your unit tests won’t enumerate. We wouldn’t ship either without a mode that executes the reasoning and discards the write.

The consequence is architectural. Every state boundary an agent can cross, whether a database, a message queue, or an outbound email, needs an interceptor with a shadow implementation behind the same interface. The database is only the first and most obvious one. Teams that build those seams now can upgrade models on their own schedule. Teams that skip them will upgrade when a vendor retires a model, and they’ll find out what changed in production.

Build the shadow mode before the first canary. The alternative is explaining a hallucinated database cascade to your stakeholders afterwards.