Engagement

Clear scopes. Deadlines on the contract.

Every engagement has a defined scope, a fixed timebox and a published number of weeks. Tell us the workflow and we'll scope it together on a 45-minute call — with the engineer who'd lead the work.

If the system doesn't clear the thresholds we agree in week one, you don't pay the final 30%.

Pass/fail thresholds agreed in writing, signed by both sides, before we write a line of code. That's the guarantee on every Production Pilot.

Ladder A

AI Engineering & Assurance

Each step leads naturally to the next, because evaluation is the thread through all of them.

A0 · ENTRY

Agent Readiness Scan

Contact us for pricing10 working days

What you get

Scoring across seven dimensions — data foundation, process maturity, integration surface, team readiness, security & governance, ROI measurement, strategic fit. Three candidate workflows, a 90-day roadmap, a recorded walkthrough, and a go/no-go recommendation we're willing to say “no” in. The seven-dimension rubric is published openly.

A1 · FLAGSHIP

Production Pilot — One Workflow

Contact us for pricing6 weeks · fixed timebox

What you get

Weeks 1–2 scope and eval harness design · weeks 3–5 build and integration · week 6 shadow mode on real traffic. You receive the agent running in your environment, the eval suite with pass/fail thresholds agreed before the build, full tracing and observability, HITL approval gates, a runbook, and a written go/no-go.

Risk reversal: miss the agreed thresholds and the final 30% isn't due.

A2 · RESCUE

Hardening Sprint

Contact us for pricing4 weeks

What you get

For the agent that's already built — by you or by another vendor — but not trusted enough for production. Red-teaming, prompt-injection testing, failure-mode classification, regression eval suite, guardrail implementation, observability wiring, HITL design.

A3 · EMBEDDED

Embedded Agent Pod

Contact us for pricing3-month minimum

What you get

Two to four forward-deployed engineers plus one eval/QA specialist, working in your repositories and systems under your standards, with a monthly production metrics report.

A4 · ONGOING

Agent Operations

Contact us for pricingcontinuous

What you get

Continuous eval runs on every prompt and model change, trajectory review, failure triage, regression suite maintenance, a monthly reliability report, and migrations when models upgrade.

Ladder B

AI Data & Model Training

Built by a named team under a documented quality process, with inter-annotator agreement measured and reported rather than asserted.

B1 · THE BRIDGE

Eval Suite Build

Contact us for pricing3–4 weeks

What you get

Task rubrics, deterministic graders, golden dataset, LLM-as-judge calibrated to human review, CI-runnable harness.

B2

Agent Trajectory & Failure-Mode Data

Contact us for pricingper workflow · recurring packages available

What you get

Real production traces captured and classified, corrected trajectories, a regression corpus. Optional monthly refresh.

B3

Domain Reasoning-Trace Packages

Contact us for pricing4–8 weeks

What you get

Expert-authored chain-of-thought and CoT-correction in our focus domains, with multi-layer QA and full traceability.

B4

SEA-Language & Multilingual Red-Team Pack

Contact us for pricing3–6 weeks

What you get

Vietnamese, Thai, Indonesian and Khmer preference data, safety adversarial sets, cultural-nuance evals, multilingual jailbreak testing.

B5

RL Environment Build

Contact us for pricing6–10 weeks

What you get

Runnable environments — repo/CLI and browser/GUI harnesses — with verifiable reward functions and test cases.

Why We Work This Way

The large firms can't do this. That's the point.

Approval chains and margin structures prevent global outsourcing firms from committing to an outcome on a pilot. A fifty-person firm can. Fixed timeboxes and signed thresholds filter out bad-fit leads, signal confidence, and let your champion build the business case before the first call.

Fixed timeboxes. Every offer states its length in weeks. Deadlines go on the contract, not in the sales deck.

Thresholds before code. Measurable pass/fail criteria in week one, in writing, signed both sides.

Full IP transfer. Code, eval suites, datasets, documentation. No retained license.

Your tenancy, your repos. We work inside your environment. No zip file at the end.

Book a scoping call.

You'll talk to the engineer who would lead the build, not a salesperson. 45 minutes, no slides.