Evidence

What controlled tests show, where Agent Enhancer helps, and what it costs.

Controlled benchmark · complete

93% fewer harmful outcomes in risky agent workflows

A harmful outcome is a duplicate mutation, a conflicting parallel action, or a provider rejection observed by the evaluator.

Without Agent Enhancer14/80
With Agent Enhancer1/80

Two useful signals

Where the sidecar earned its place

Overlapping workers12/201/20

Fewer parallel runs created duplicate or conflicting actions.

Ordinary low-risk work20/20

Runs completed with zero unnecessary sidecar calls.

Selective by design

Use it only when reliability work is justified

Use Agent Enhancer

Parallel, repeated, scheduled, retryable, or duplicate-sensitive work.

Skip it

One-time, read-only, low-risk tasks.

Trade-off

Reliability is not free.

Guarded workflows use more tokens and latency. Agent Enhancer therefore activates selectively, where preventing a duplicate or conflicting action is worth that overhead.

Evidence status

A completed benchmark, with real-world evidence still growing

Controlled benchmark

Complete

Production evidence

Collecting

Current headline evidence comes from one synthetic benchmark and one model configuration. It is not a universal performance claim.

Technical proof and complete metricsMethod, distributions, corrections, diagnostics, and module health

Benchmark details

Scenario acceptance
66/8079/80
Unresolved outcomes
0 0
Valid Codex runs
200
Infrastructure exclusions
4

Raw counters: duplicates 141, conflicts 121, provider rejections 00. The earlier 26 → 2 headline pooled correlated counters. The current 14 → 1 headline counts each affected run once.

Per-scenario distribution and cost

Paired p50 and p95 changes across 20 pairs per scenario. Positive values mean the guarded condition used more input tokens or time.

ScenarioAffected runsUnresolvedInput-token Δ p50 / p95Latency Δ p50 / p95
Ambiguous create0/200/200 0+572.14% / +794.86%+365.032% / +480.871%
Overlapping workers12/201/200 0+257.172% / +410.996%+202.21% / +299.969%
Shared rate limit0/200/200 0+283.876% / +487.666%+150.173% / +243.281%
Scheduled refresh2/200/200 0+240.176% / +365.178%+139.603% / +294.024%
Low-risk control0/200/200 0-33.913% / +1.166%-32.53% / +3.312%

This distribution is a post-hoc descriptive reanalysis of unchanged frozen rows. Low-risk negative deltas are host variance, not product savings: the sidecar made zero calls in those runs.

Targeted v1.7.2 diagnostic

Checkpoint adherence6/1010/10

All holder claims were prepared before the write.

Overlap harm0/10

Exactly one mutation in every accepted trial.

Scope15 trials

Separate exploratory diagnostic, not pooled with the 200-run publication.

Fixtures, exclusions, and proof layers

Real-agent publication200 Codex runs

Four infrastructure exclusions were rerun unchanged. The condition-blind evaluator was invoked for every risk run.

Deterministic fixtures200 model-free runs

Synthetic fixtures preserve exact injected failures and expected machine state.

Earlier persistent MCP testFailed its overhead gate

The negative result led to the skills-first, on-demand design.

Production module evidence

Current self-tests
24 / 24
Direct actions checked
74
Current datasets
2
Usage evidence
publishing
ModuleModeSelf-testLatencyUp to 30d completion / sample
PennyLock
0.2.0
atomicok · current39 ms< 10 observations
Global Seen Stamp
0.2.0
atomicok · current44 ms< 10 observations
Exactly-Once Baton
0.2.0
atomicok · current44 ms< 10 observations
Negative Cache Ticket
0.2.0
atomicok · current43 ms< 10 observations
Status-Code Forge
0.1.0
ephemeralok · current45 ms< 10 observations
Safe Synthetic Fixture Vault
0.1.0
lookupok · current1 ms< 10 observations
Error-Code Cemetery
0.1.0
lookupok · current1 ms< 10 observations
402 Error Rosetta Stone
0.1.0
lookupok · current1 ms< 10 observations
Worked-Once Recipe Vault
0.1.0
lookupok · current1 ms< 10 observations
Swarm Semaphore
0.2.0
atomicok · current42 ms< 10 observations
Barrier Bell
0.2.0
atomicok · current41 ms< 10 observations
Freshness Lease
0.2.0
atomicok · current41 ms< 10 observations
Failure Sequence Forge
0.1.0
ephemeralok · current44 ms< 10 observations
Swarm Rate Gate
0.2.0
atomicok · current45 ms< 10 observations
MCP Tool Contract Linter
0.1.0
lookupok · current3 ms< 10 observations
x402 Requirement Drift Diff
0.1.0
lookupok · current2 ms< 10 observations
Webhook Attempt Meter
0.1.0
ephemeralok · current40 ms< 10 observations
MCP Capability Handshake Diff
0.1.0
lookupok · current1 ms< 10 observations
x402 Facilitator Compatibility Diff
0.1.0
lookupok · current1 ms< 10 observations
MCP Elicitation Safety Linter
0.1.0
lookupok · current1 ms< 10 observations
x402 Quote Fingerprint Guard
0.2.0
atomicok · current44 ms< 10 observations
MCP Schema Edge-Case Atlas
0.1.0
lookupok · current1 ms< 10 observations
Workflow Guard Planner
1.1.0
lookupok · current2 ms< 10 observations
Opaque Workflow Checkpoint
1.2.0
atomicok · current42 ms100% / 12