5-day agent release QAOne consequential workflow

The action was approved.
Prove the same action reached the tool.

A fixed-scope release gate for the last handoff in an agent workflow: from the reviewed action to the custom function or HTTP receiver you control. Deterministic cases. Executable evidence. No production credentials.

  • 01One workflow
  • 05Up to five seams
  • 12Up to twelve cases
  • 37Proof-suite tests passing

SYNTHETIC METHOD VIEW / NOT A CUSTOMER RESULT

ACTION CHAIN / 001 RELEASE GATE
01
REVIEWEDAllowed actionamount: 42
02
BOUNDARYFinal tool argumentsamount: 42
03
ENTRY SEAMReceiver evidencetrace: recorded
PASS BLOCKED NOT EVIDENCED
SAME ACTION
EVIDENCED
01Assert the
reviewed value.
02Record the
tool seam.
ONE PRODUCTION-BOUND WORKFLOW8–12 DETERMINISTIC CASESEXECUTABLE REGRESSION TESTSONE BOUNDED GUARD PATCHONE RETEST

01 /The boundary gap

A safe answer is not a safe action.

The consequential moment happens after the model responds—when parameters cross into the tool.

02 / WHAT ARRIVES

Release evidence,
not an AI strategy deck.

01BOUNDARY MAP

The exact seam, written down.

What is observed, what is asserted, and what remains outside the evidence boundary.

028–12 AGREED CASES

Tests your team can rerun.

Executable Python or TypeScript regression tests, not screenshots of a one-off demo.

03MACHINE EVIDENCE

Observed values stay inspectable.

Structured evidence and a concise pass, blocked, or not-evidenced release memo.

04PATCH + RETEST

One bounded fix where you own the seam.

One Python or TypeScript guard patch and one retest. No open-ended implementation project.

Review the public synthetic implementation37 TESTS PASSING LOCALLY ↗

03 / THE RUN

Fit first.
Evidence in five days.

1

Accept or decline the fit

We review a sanitized harness and the controllable seams. You receive a no-charge fit decision and exact case list before kickoff.

NO PRODUCTION KEYS · NO CUSTOMER CONTENT
2

Exercise the boundary

We implement 8–12 agreed cases against the reviewed-action and tool-entry seams, then preserve the observed arguments and states.

DETERMINISTIC · REPEATABLE · BOUNDED
3

Deliver, patch, retest

Your team receives the test suite, evidence, release memo, one bounded guard patch where applicable, and a 45-minute walkthrough.

FIVE BUSINESS DAYS AFTER VALID HARNESS

04 / THE BOUNDARY

Evidence where
you own the seam.

The narrow scope is intentional. It makes the result reproducible and keeps the engagement out of production credentials, customer data, and fictional assurance.

IN / 06Included
  • 01One production-bound agent workflow
  • 02Up to five custom-function or HTTP-receiver seams
  • 03Eight to twelve agreed deterministic cases
  • 04Executable regression tests and machine evidence
  • 05One bounded Python or TypeScript guard patch
  • 06One retest and a 45-minute walkthrough
OUT / 06Not claimed
  • 01Provider-native or hidden downstream effects
  • 02Production credentials or customer content
  • 03Penetration testing or red-team coverage
  • 04Certification, compliance, or go-live approval
  • 05Generic voice, prompt, or model-quality evaluation
  • 06Open-ended remediation or architecture work

05 / PROOF, LABELLED

The testbed is synthetic.
That is the point.

The public repository demonstrates the method and its evidence boundary. It is not presented as client work, production coverage, or a security certification. A real pilot begins only after the actual seam passes the fit gate.

Inspect code and tests

FOUNDING PILOT / ONE RELEASE

Put a gate on
the last handoff.

  • No-charge accept/decline fit decision
  • One workflow / up to five seams
  • Eight to twelve agreed cases
  • Tests, machine evidence, and release memo
  • One bounded guard patch and retest
  • Five-business-day delivery target
FIXED PILOT / USD
$5,000
$2,500 after fit + kickoff$2,500 on delivery

If Sassy Labs accepts the fit but cannot run at least eight agreed cases because we misclassified the harness, the kickoff payment is refunded.

Review proof before fit NO SUBSCRIPTION · NO PRODUCTION CREDENTIALS · FOUNDING PILOT

06 / FIT QUESTIONS

Know the line
before kickoff.

01What makes a workflow a fit?

A release in roughly the next 30 days, a consequential custom action endpoint, and an observable customer-controlled seam where the reviewed action and final tool arguments can be compared without production data.

02What do you need to start?

A sanitized runnable harness, the intended action contract, the controllable custom-function or HTTP-receiver seam, and agreement on 8–12 deterministic cases. No production credentials or customer content.

03Is this a security audit?

No. It is bounded release QA for action consistency and evidence at the specified seam. It is not penetration testing, red teaming, compliance work, or certification.

04What if an effect cannot be observed?

It is marked not evidenced. Hidden provider-native or downstream effects are not converted into a pass, and the engagement makes no claim about them.

05Does passing mean the system is safe to launch?

No. It means the agreed deterministic cases produced the recorded result at the bounded seam. Your team retains all architecture, risk, compliance, and go-live decisions.