Production Triage
The smallest engagement that can reduce your uncertainty. Most relationships start here.
- Architecture review — how your system actually works, verified from code, including where your docs disagree with reality. PDF + editable diagram
- Prioritized risk map — every material risk, ranked by blast radius, each tied to a specific component. PDF
- Remediation plan — the top risks as bounded work items with acceptance criteria any senior engineer could execute. PDF
- 45-minute walkthrough — live with your team, recorded if you want it, ending in a clear recommendation: build, audit deeper, or stop. call + recording
Done means: you can name your highest-leverage technical move and defend it to your own team without me in the room. Fee credits toward a sprint started within 30 days.
Production Audit
For systems where the risk surface needs more than three days, or where you want specs your team can execute alone.
- Everything in Triage, at full depth.
- Code-level findings — annotated references into your repo covering contracts, data flow, failure modes, release discipline. written report
- Implementation-ready specs — acceptance criteria, migration and rollback notes, test expectations per work item. one spec each
- AI evaluation notes (when in scope) — how your prompts, retrieval, and tool use behave under real inputs, with an eval harness you keep. report + scripts
- Sequenced roadmap — now / next / later with dependencies and rough cost. 1 page
Done means: your team could start any recommended item tomorrow morning without a clarifying meeting.
Build or Rescue Sprint
Hands-on delivery against one bounded production outcome.
- Working, deployed code in your repo, under your CI, through reviewed pull requests — never a zip file. PRs in your repo
- Tests that guard the outcome — the behavior you bought is the behavior that's tested. suite in CI
- Weekly written status — shipped, in progress, blocked, decisions needed. Short enough to actually read. async, your channel
- Decision log — every non-trivial choice recorded with context, so the reasoning outlives the engagement. markdown in repo
- Handoff package — runbook with deploy, rollback, and known failure modes; updated diagram; recorded walkthrough. docs + video
Done means: the outcome is live, your team can operate and extend it, and nothing depends on me being reachable.
Fractional Engineering
Ongoing senior ownership without the full-time hire.
- A named owner for an agreed slice of your system — including its incidents and its roadmap. 1-page charter
- Shipped work every month against priorities we set together, under the same PR-and-tests discipline as a sprint. PRs + demo
- Monthly engineering report — what shipped, system health, accumulating risks, and what I'd fund next if it were my money. 1–2 pages
- Continuity guarantee — decision log and runbooks stay current, so offboarding is a two-week handoff, not a crisis. living docs
Done means, every month: the report is sent, the demo happened, and the docs reflect reality.
Standards on every engagement
These aren't add-ons; they're how "AI systems that hold up" gets kept as a promise.
Work happens in your house
Your repo, your cloud, your secrets. You lose nothing if the engagement ends tomorrow.
Written over verbal
Every recommendation and decision exists as a document you keep. Calls explain documents; they don't replace them.
Honest exits
Every engagement can end with "stop — this isn't worth building." A clean no is a deliverable.
Failure modes named up front
Anything shipped comes with its limits and rollback path documented.
The path in
A 15-minute fit call establishes the problem, urgency, and decision owner. If there's a fit, Production Triage finds your highest-leverage move in three days — and the fee credits toward a sprint. See a real example: sample triage report.
Book a fit call