Your board asked what you're doing about AI. Eight months later there are pilots in four departments, a licence bill nobody can fully explain, and a slide deck that says “promising.” Ask a simple question — which of these is actually working, and who owns it? — and the room goes quiet.
That's not an AI problem. It's an operating problem, and it has a shape you already recognise.
Sound familiar?
- Pilots started in four departments, and no two are measured the same way.
- You bought 200 seats. You have no idea how many people used one last week.
- Every initiative has a champion. None has an owner — the person whose name is on the outcome.
- Spend arrives on the card monthly and gets discussed annually.
- Nothing has ever been killed. Pilots don't end; they just get quieter.
- The people doing the work have opinions about what's working. Nobody has asked them in a structured way.
Why AI programmes stall
AI doesn't fail at the pilot. It fails at the habit. The demo works — that's the easy part now. What breaks is the sixty days after, when using the new tool is still slower than the old way and nobody's checking whether anyone made the switch. Adoption is a behaviour change problem wearing a technology costume.
Status decks are lagging indicators. By the time a monthly update says a pilot is struggling, you've spent a month of licence fees and burned the enthusiasm of the people who volunteered. The interval between “this isn't working” being true and being said out loud is where the money goes.
Committees own tools; nobody owns outcomes. An AI steering group with nine members and a shared mandate produces meeting notes. One named person accountable for one number produces movement. This is the least popular sentence in this blueprint and the most reliable one.
The blueprint
Six pieces. None of them are novel — that's the point. This is the operating rhythm good companies already run on their core business, pointed at the AI portfolio.
One line on what AI is for here
Not a strategy document. One sentence everyone can repeat: “We're using AI to cut the time it takes to answer a customer, without making the answer worse.” Every initiative either serves that line or gets asked why it exists.
Three to five bets, one owner each
Not fifteen experiments. Each bet gets a single named owner — a person who runs that function, not someone from IT — a start date, and a date by which you'll decide its fate. Write the kill criteria at the start, while you're still calm about it.
A weekly scorecard
Eight to twelve numbers, same numbers every week, each with an owner and a target. Adoption, quality, outcome, spend. It should take 90 seconds to read and it should occasionally make someone uncomfortable.
An issues list that stays visible
“Legal hasn't cleared the vendor.” “Sales won't use it because it writes like a robot.” Write them down, give each one an owner, and let unsolved ones visibly carry week to week. Carried issues are the most honest metric in the building.
Thirty minutes, every week, same time
Not a project update. A working meeting where the numbers are read, the owners speak, the top issues get solved, and everyone leaves with a commitment. Weekly is the whole trick: it's short enough to sustain and frequent enough to catch drift before it costs a quarter.
A quarterly decision, not a review
Every 13 weeks each bet gets scaled, killed, or explicitly extended with a new decision date. “Still going” is not a status. A portfolio that has never killed anything isn't a portfolio — it's a collection.
What to actually measure
The scorecard is where most AI programmes go wrong in one of two directions: a dashboard of forty model metrics nobody reads, or a single vanity number (“10,000 prompts this month!”) that proves nothing. Pick eight to twelve across these four categories, and make sure at least one of them can embarrass you.
| Metric | Why it earns its row | Usual owner |
|---|---|---|
| Adoption — is anyone actually using it? | ||
| Weekly active users, as % of licensed seats | The most honest number on the page. Seats bought is not seats used, and the gap is usually a shock the first time you look. | Function owner |
| Seats with zero use in 30 days | Money you are actively wasting — and a list of people worth talking to about why. | Function owner |
| Depth: users running 5+ assisted tasks a week | Separates real adoption from people who logged in once and left. Usually about a third of your "active" number. | Function owner |
| Teams with at least one live use case | Spread. Catches the programme that looks healthy because one enthusiastic department is carrying it. | Exec sponsor |
| Quality — is it any good? | ||
| Human override / correction rate | Trust, measured. If your team rewrites every output, you have a demo, not a tool. Watch this fall as prompts and training improve. | Process owner |
| Escalations from AI-handled work | The cost of being wrong. Deflection numbers mean nothing without this beside them. | Support lead |
| Output accepted without an edit | The optimistic twin of override rate, and easier to instrument in some tools. Track one or the other, not both. | Process owner |
| Rework hours caused by AI output | Makes the hidden tax visible next to the claimed savings. The number most programmes never look at. | Process owner |
| Outcome — did the thing you wanted actually happen? | ||
| Cycle time on the target process | The reason you started. If this doesn't move, the rest is theatre — however good the adoption looks. | Process owner |
| Tickets deflected, or throughput per person | Capacity change in a number you can put in front of a board. | Function owner |
| Quality score on the target process | The guardrail. Proves the speed didn't come out of the quality — without it, cycle time is a half-truth. | Process owner |
| Estimated hours returned to the team | Self-reported and imperfect. Useful as a trend you watch, dangerous as a claim you publish. | Function owner |
| Spend & portfolio — what is this costing, and are we deciding? | ||
| Total AI spend this period | Tools, tokens, and consultants together, or the number lies. Monthly is too slow to catch a runaway. | Finance / sponsor |
| Cost per assisted task | The unit economic. Falling means leverage; rising means a pilot worth questioning. | Finance / sponsor |
| Initiatives in flight vs. graduated vs. killed | Focus and decisiveness in three numbers. Most companies are heavy on the first and empty on the third. | Exec sponsor |
| Spend with no named owner | Shadow AI, in dollars. Usually the most uncomfortable row on the page — and the fastest saving available. | Finance / sponsor |
| Days since the oldest pilot started | Catches the pilot that quietly became permanent without anyone ever deciding it should be. | Exec sponsor |
The starter pack includes a 26-metric library with definitions and suggested owners, so you can choose rather than invent. Choosing eight is the work. Adding thirty is the avoidance of it.
The 30-minute meeting
Who's in the room
What good looks like at 90 days
What this blueprint isn't
It isn't AI governance. It won't monitor models, log prompts, evaluate bias, catalogue training data, or satisfy an auditor — and any page that claims one framework does all of that is selling something.
It also isn't a maturity model or a transformation programme. It's a meeting, a scorecard, and the discipline to hold both for thirteen weeks. Pair it with whatever compliance tooling your risk function requires.
This is a blueprint, not a case study. We haven't published customer results here because we haven't earned them yet — when we have, they'll appear with names attached. Everything above comes from twenty years of watching operating rhythms succeed and fail in small companies.
Running it in a spreadsheet, or running it in Vetta
The starter pack works. Plenty of good companies run this rhythm on a spreadsheet for a year, and if that's where you start, you'll get most of the value.
What a spreadsheet can't do is make the meeting live: carry the unsolved issue forward automatically, show whose check-in has gone stale, keep the tasks that come out of the meeting attached to the initiative they serve, and hand you a quarter's worth of evidence when the kill-or-scale conversation arrives. That's what Vetta is — the same loop, with the meeting as the product rather than a calendar invite.
Notice what the screen makes unavoidable: the killed initiative is still visible with the reason it died, the override-rate issue is carrying its fourth week in red, and one bet is past its decision date — so Thursday's meeting has to deal with it. That's the whole mechanism. Nothing here is AI-specific; it's just an operating rhythm that refuses to let things go quiet.
One note that matters for this particular use case: Vetta is priced flat, per company, not per seat. Your AI council is eight people — but your whole company is already included. When the rhythm proves itself with the council, rolling it out to everyone else costs nothing extra. That's unusual, and it's deliberate: accountability shouldn't have a meter on it.