Track everything, silently
Bolt a monitoring agent onto every IDE and you get compliance-shaped data: engineers route around it, sanitize their prompts, or quietly stop using AI where it's watched. You end up measuring fear, not skill.
Kata isn't a one-time audit — it's a repeatable exercise your engineers run each quarter, so improvement shows up as a trend, not a guess. Anonymized by default: you see the spread and the trend, never a dossier on one person.
Scoreboard above is illustrative — each row compares against that participant's own last run.
AI usage inside an engineering org is either invisible or surveilled — and both leave you guessing about where your AI budget is actually going.
Bolt a monitoring agent onto every IDE and you get compliance-shaped data: engineers route around it, sanitize their prompts, or quietly stop using AI where it's watched. You end up measuring fear, not skill.
You're paying for Copilot, Cursor, and Claude seats across the whole org with no idea which teams are getting real leverage from them and which are burning tokens on trial and error.
There's a third option: an anonymized, opt-in practice kata, run quarterly so the goal is visible improvement, not a permanent record. Cohort-level insight instead of a dossier on each engineer — identity revealed only when someone chooses to be.
See how it worksSame replay-driven engine as hiring, run on a cadence so you see improvement, not just a single snapshot.
Pick a standardized task (or bring your own), invite a cohort — a team, a business unit, or the whole company — and set the budget and window.
Everyone is graded and judged the same way hiring candidates are — but results default to anonymous. The spread and the patterns are visible; identities aren't, unless someone opts in.
A coaching report compares top and bottom results — prompting discipline, model choice, context handling — and turns it into concrete advice, not a leaderboard to hide from.
Each quarter's kata lands next to the last one, so you can see whether prompting discipline, model choice, and cost per outcome are actually improving — not just where things stand today.
Same replay-driven signal as hiring, pointed at your own team's day-to-day AI workflow.
Results are pooled and ranked without names. Reveal is opt-in — top performers can choose to be credited; nobody is exposed for scoring low.
Kata runs on a cadence, so this quarter's result sits next to last quarter's. The point is the trend, not a single snapshot — and it shows whether the coaching is actually landing.
The event log and LLM judge are the same ones already proven on candidate assessments — not a separate, unproven tracking tool bolted on afterward.
See who's escalating to expensive models out of habit versus need, and who's getting more done on cheaper ones — the real driver behind your AI spend.
Opt-in top runs become a growing internal library your team can learn from — a compounding asset, not a one-off report.
Use the standard kata catalog, or author tasks against your own codebase and patterns.
We're onboarding teams by hand while we learn what works. Join the wait list and tell us the size of your org — we'll get a cohort set up.