
Engagement
Clear scopes. Deadlines on the contract.
Every engagement has a defined scope, a fixed timebox and a published number of weeks. Tell us the workflow and we'll scope it together on a 45-minute call — with the engineer who'd lead the work.
If the system doesn't clear the thresholds we agree in week one, you don't pay the final 30%.
Pass/fail thresholds agreed in writing, signed by both sides, before we write a line of code. That's the guarantee on every Production Pilot.
Ladder A
AI Engineering & Assurance
Each step leads naturally to the next, because evaluation is the thread through all of them.
What you get
Scoring across seven dimensions — data foundation, process maturity, integration surface, team readiness, security & governance, ROI measurement, strategic fit. Three candidate workflows, a 90-day roadmap, a recorded walkthrough, and a go/no-go recommendation we're willing to say “no” in. The seven-dimension rubric is published openly.
What you get
Weeks 1–2 scope and eval harness design · weeks 3–5 build and integration · week 6 shadow mode on real traffic. You receive the agent running in your environment, the eval suite with pass/fail thresholds agreed before the build, full tracing and observability, HITL approval gates, a runbook, and a written go/no-go.
Risk reversal: miss the agreed thresholds and the final 30% isn't due.
What you get
For the agent that's already built — by you or by another vendor — but not trusted enough for production. Red-teaming, prompt-injection testing, failure-mode classification, regression eval suite, guardrail implementation, observability wiring, HITL design.
What you get
Two to four forward-deployed engineers plus one eval/QA specialist, working in your repositories and systems under your standards, with a monthly production metrics report.
What you get
Continuous eval runs on every prompt and model change, trajectory review, failure triage, regression suite maintenance, a monthly reliability report, and migrations when models upgrade.
Ladder B
AI Data & Model Training
Built by a named team under a documented quality process, with inter-annotator agreement measured and reported rather than asserted.
What you get
Task rubrics, deterministic graders, golden dataset, LLM-as-judge calibrated to human review, CI-runnable harness.
Agent Trajectory & Failure-Mode Data
Contact us for pricingper workflow · recurring packages availableWhat you get
Real production traces captured and classified, corrected trajectories, a regression corpus. Optional monthly refresh.
What you get
Expert-authored chain-of-thought and CoT-correction in our focus domains, with multi-layer QA and full traceability.
What you get
Vietnamese, Thai, Indonesian and Khmer preference data, safety adversarial sets, cultural-nuance evals, multilingual jailbreak testing.
What you get
Runnable environments — repo/CLI and browser/GUI harnesses — with verifiable reward functions and test cases.
Why We Work This Way
The large firms can't do this. That's the point.
Approval chains and margin structures prevent global outsourcing firms from committing to an outcome on a pilot. A fifty-person firm can. Fixed timeboxes and signed thresholds filter out bad-fit leads, signal confidence, and let your champion build the business case before the first call.
Fixed timeboxes. Every offer states its length in weeks. Deadlines go on the contract, not in the sales deck.
Thresholds before code. Measurable pass/fail criteria in week one, in writing, signed both sides.
Full IP transfer. Code, eval suites, datasets, documentation. No retained license.
Your tenancy, your repos. We work inside your environment. No zip file at the end.
Book a scoping call.
You'll talk to the engineer who would lead the build, not a salesperson. 45 minutes, no slides.