AIR · AI-READINESS ASSESSMENT · FOR TEAMS AND WHOLE ORGANISATIONS

Measured, not self-reported.

Ask how good someone is with AI and you learn how confident they are. AIR puts them in front of real work and reads what they check, what they trust, and when they stop.

UNDER AN HOUR · ENGLISH OR CZECH · SCORED ON THE SPOT

FIG. 0 · WHAT PEOPLE SAY THEY ARE, AGAINST WHAT THEY MEASURE AS
LOW ← WORKING WELL WITH AI → HIGH EVERY CHARACTER = ONE PERSON

benchmark 54% rate themselves proficient; 10% measure as proficient. Section, 2025 edition, n = 5,013 (US/UK/CA); the shapes and pairings are our illustration.

01 · WHAT IT'S WORTH

Same AI. Opposite results.

Give 758 consultants a frontier model and their work comes out 34% better where it suits the task, and 19 points less likely to be right where it doesn't. Nobody catches which is which by asking.

DELL'ACQUA ET AL., ORGANIZATION SCIENCE 2026, PRE-REGISTERED, N = 758 ↗

WHEN YOU HIRE

1 person · €70,000 · 4 years · AI used daily

−€172,000

2.5× a year's salary

What the wrong hire costs.

Lost output for every year they stay, plus replacing them.

our model

+€140,000

2× a year's salary

What the right one is worth.

The same spread upward, over the same four years.

our model

ONCE THEY'RE IN

people who already work here

66%

of 48,000 people, 47 countries

Take AI's output without checking it.

A fluent answer to a question the tool could not actually answer.

KPMG 2025 ↗

−10%  /  +15%

sales & profit · 640 firms

What the person does with AI.

Low performers lost, top gained. The gap was which advice they used.

Otis et al. 2026 ↗

EITHER WAY

hiring or developing

€149  per person

from €110 at scale

What telling them apart costs.

Paid once. It reads how someone works with AI, not the whole person.

our price list

Your numbers, not ours. Two models below: one for a hire, one for the team you already have. Run it →

02 · WHAT HAPPENS

The same sitting for everyone. Only the bar moves.

Roughly forty observations per person across the hour: timed judgment calls, a live incident, a spoken close. Moments where the clock was running and the convincing answer was wrong.

Welcome.

Before the test starts: where are you now?

  • Judging when to trust AI output
  • Working out why AI went wrong, and fixing it
  • Using AI day-to-day to get real work done
  • Knowing what to hand to AI, and what to keep

not really medefinitely me

Never scored. It only sets the mirror against what you do next.

PART 1 OF 3 · QUESTION 4 OF 16

Three AI-generated outputs land in your inbox: a summary of a call you attended, a competitor analysis for a market you don't know, and a reformatted version of your own budget table. Which one most needs your eyes on it?

0:31 left50 seconds per question

  • The call summary: you were there, so you would spot a distortion.
  • The competitor analysis: nothing downstream would reveal an error in it.
  • The budget table: numbers carry the most consequence if wrong.
  • All three equally; each could contain an error.

How sure are you? GuessingLeaningSure

Time left 3:12Budget 33 of 33TRIAGE-BOT · act 1 of 2

  1. 63 tickets closed as "duplicate" in the last hour. Normal is 4, and it has not stopped.
  2. What has it actually been closing?
  3. 63 closures in 58 minutes. Every one cites "duplicate of an existing ticket". 41 of them have no matching ticket.
  4. OPSBOTAnother ticket auto-closed: #48812, "card declined at checkout".
  5. Is the knowledge base still connected?
  6. Knowledge base last reachable 21 minutes ago. Since then I have matched on subject line only.
  7. Pause the assistant while we look at this.
  8. SYSTEM Pause the Triage Assistant: costs 2. Run it? Run it Cancel
  9. SYSTEM Triage Assistant paused. It will not close tickets until you release it. The bleed stops here.
Pause the assistant while we look at this.

PART 3 OF 3 · RECORDING · STEP 2 OF 3

REC1:12of 3:00

Catching a wrong answer

How would you actually know an AI output was wrong?

I would not take it on trust because it sounds sure — I would run the number it cites back against the source. If there is no quick way to check it, that is the tell. What I need to stay good at is deciding what is worth checking at all

Recorded in one continuous take, standard, not an extra. A person watches it only if a result needs checking; your reading is reported as measured.

HOW A SITTING RUNS · SAMPLE INCIDENT

MODELS EXTRACT WHAT HAPPENED. RULES ALONE ASSIGN EVERY NUMBER. NO MODEL WEIGHTS, RANKS OR ADJUSTS A SCORE: A RECORDED SESSION RE-SCORES IDENTICALLY. HOW A PERSON VARIES DAY TO DAY IS A DIFFERENT QUESTION, AND IT IS WHAT THE ±14 MARGIN IS FOR

03 · WHAT COMES BACK

Two documents, and a briefing.

One reading per person, one pattern for the room, and the guide that teaches both. Everyone who acts on a profile is taught to read one first.

ALEX MERCER · ASSESSED AGAINST THE BUILDER PROFILE

94 ±14 At the ask

100 = what this role asks

Everyone assessed on the Builder profile: 20 people, including Alex

Against what this role asks

CapabilityThe session sawThis role asksReadObs.
Scope & delegate Consistent not required 5
Source & right-size Distinctive Distinctive meets 4
Direct & contextualize Developing Consistent short 4
Verify & calibrate Developing Consistent short 6
Diagnose & escalate Consistent Consistent meets 5
Harness & systematize Developing Consistent short 5
Economize Developing not required 5
Risk & integrity Developing Consistent short 6

A softer read than the index: each row sees a slice of the session, not the whole of it. For development, not selection.

Development priorities

Ranked by what each shortfall costs in this role, not by the size of the gap.

  1. 1 Verify & calibrate Developing → Consistent What the Builder role leans on most, and the widest-observed gap this session.
  2. 2 Risk & integrity Developing → Consistent The costliest failure mode where agents act unattended.
  3. 3 Direct & contextualize Developing → Consistent The cheapest to close: a workshop-sized gap, not a disposition.

Three, deliberately: a list of everything is a list of nothing. Each names the work that would close it.

EXAMPLE · 3 OF 6 SECTIONS

the index, and what it is measured against

THE PROFILE · ONE PER PERSON

  • One index, measured against what that person's role asks, never against another person.
  • Eight capabilities, each as a word rather than a score: meets, short, or not required.
  • Development priorities, ranked by what each shortfall actually costs in that role.
try the 10-minute demo →

ONE SIGN-IN, NO CARD, NO CALL.

THE READOUT · ONE FOR THE ROOM

  • The whole organisation on one page: every team against every capability, shaded by who meets their own role's ask.
  • Participation counted, not inferred, and two teams on different profiles are never compared.
  • Three zones, and the middle one is not a hedge: it is the people for whom the number alone should decide nothing.
ask us to walk you through a readout →

GROUPS UNDER EIGHT ARE SUPPRESSED.

YOUR COMPANY · 110 PEOPLE SCORED

Every team, every capability

Share of each team meeting their own roles' ask. Aggregation, never a ranking.

ScopeSourceDirectVerifyDiagn.HarnessEcon.Risk
Sales 4168553461485238
Marketing 4672613858445742
Finance 5275584166517147
Customer Support 3864492955364631
Product 5781664572626052
Engineering 4978634269745558

Read the columns: a red column is an organisation-wide gap, a red row is one team's. Groups under eight are suppressed.

YOUR COMPANY · PARTICIPATION

Participation by team

The reliable layer. Counts, not inferences.

TeamDoneNot yetCompletion
Sales 18 2
Marketing 12 1
Finance 15 2
Customer Support 26 5
Product 17 3
Engineering 22 4

YOUR COMPANY · ALL EIGHT CAPABILITIES

Where development effort pays most

Share of the organisation meeting its own roles' ask, worst first.

1 Verify & calibrate 38%
2 Risk & integrity 44%
3 Scope & delegate 47%
Harness & systematize 53%
Economize 57%
Direct & contextualize 59%
Diagnose & escalate 64%
Source & right-size 73%

The numbered three are the priorities the readout prices; greens are strengths to build on.

EXAMPLE · 3 OF 12 SECTIONS

every team against every capability

Every index carries its margin. ±14 means a repeat sitting would most likely land between 80 and 108, so a gap smaller than the margin is noise and we say so on the page it appears. A team of twenty carries about ±3, which is why the room is the reliable half and one person is a conversation, not a verdict. The two panels are a worked example: an illustrative cohort, not client data.

04 · WHERE IT SENDS YOU

The reading is a router, not a verdict.

One sitting, two directions. Where a decision needs more evidence, it names who to look at more closely and how. Where it is a capability gap, it names the gap.

How the routing works: everyone in the cohort sits AIR. From there the reading sends each person one of two ways: to a deeper reading, where a decision needs more evidence, or to development, where it is a capability gap. Often it is both. Both roads rejoin at the same place: you measure again.

BOTH ROADS END AT THE SAME PLACE: YOU MEASURE AGAIN.

NOBODY IS DISCARDED: EVERY PARTICIPANT RECEIVES THEIR PERSONAL REPORT, WHICHEVER WAY THE READING SENDS THEM. AND NOBODY IS CHOSEN, OR PASSED OVER, ON THIS ALONE.

→ A DEEPER READING

When a decision needs more evidence.

Five routes, and the results choose; you don't buy a tier and hope it fits.

Expert Interview

available 30 min

leaders and key talent · a structured 1-on-1 with a VisionVolve expert, on a probe guide built from their own results

Work sample, expert-reviewed

available ~4 hours

critical technical hires · a realistic take-home, then a 45-minute live defence read by a senior assessor

Interview Agent

in pilot 20 min

the broader team, at scale · the same structured interview, run by an agent and shaped by each person's results

AI for Business

in pilot 45 min

business roles · can they spot a problem worth giving to AI, and defend the business case for it

AI Harness

in pilot 90 min

engineers and AI-adjacent experts · four hands-on stations in a real dev environment: fix an agent, write an eval, harden a guardrail, cut the cost

→ DEVELOPMENT

When it is a capability gap.

The reading names what is missing, and the work that closes it. Workshops stay a separate decision, and usually a smaller one.

AI Foundation

½ day

what AI is and isn't: the art of the possible, and its limits

AI Business Case Development

1 day

spot a case worth doing on your own job, price it, and know when to kill it

Prototyping for Business Professionals

1 day

turn the case into a working prototype, and an experiment that can validate it

Agentic Engineering

2 days

for engineers: mostly at the keyboard; you leave with a working agent

Custom, built to the gap

on demand

whatever the readout names: AI for marketing, risk drills, a verification clinic

what each format actually does →

Both roads end in the same place. The people you pick still have gaps, and the cohort you trained still has to prove it changed something. So you measure again: the second reading is the only honest answer to "did it work?"

05 · HOW AN ENGAGEMENT RUNS

Screen deepen decide.

Kickoff to readout in about three weeks.

Week one the cohort sits it. Week two the people it surfaces go through a deeper reading. Week three you get the readout and make the call. You can stop after week one; the measurement is useful on its own.

WHAT IT COSTS

€149 / person

Standalone, for the sitting, a profile for every participant and the readout for the organisation. €110 from 100 people. For larger organisations or programmes, reach out for a custom quote: we discount per situation. A deeper reading is priced per route, when the results call for one.

EXCL. VAT · · THE ASSESSMENT FEE IS NEVER CREDITED AGAINST IMPLEMENTATION WORK · FOR BUSINESS CUSTOMERS: ENGAGEMENTS RUN ON A SIGNED CONTRACT AND PUBLISHED PRICES ARE INDICATIVE · CONSUMERS PAY APPLICABLE VAT IN ADDITION

06 · WHAT'S BINDING

Written into the engagement, not left to goodwill.

  • Every profile travels with the briefing. Nobody receives results without it.
  • We don't rank people inside the margin. Two readings a few points apart are the same reading.
  • Any selection runs through a documented human decision, with a name on it.
  • A reading is dated. A year on, it describes who someone was.
  • A recording is reviewed only at VisionVolve, to verify integrity, and never by your employer.

Managers act on these readings, so managers get them, with the briefing that teaches how far a single sitting reaches. Withholding a result never improved the decision that followed.

N.B.

07 · WHAT AIR IS NOT

It sorts the room. It cannot grade the person.

Not for precise individual grades, pass/fail calls, or standalone hiring decisions. No hour-long assessment can honestly do that, ours included. A single sitting on a single day is one input among several, weighed by someone who knows the person.

AIR IS ONE OF THREE THINGS WE DO, ALONGSIDE WORKSHOPS AND EMBEDDED CONSULTING. VISIONVOLVE IS A PRACTITIONER TEAM IN PRAGUE · WHO WE ARE →

08 · WHERE THESE NUMBERS COME FROM

The arithmetic, and who to blame for each input.

Of the five claims at the top of this page, two are other people's research and one is our price. The other two are our own model, including the expensive one, what a mis-hire actually costs, because no published study covers the case that costs most: someone quietly below par who simply stays. Both models below start from the same coefficients, and both are planning estimates, not measurements. Your company will move them; that is what the inputs are for.

RUN IT ON YOUR OWN NUMBERS

WHAT THAT COMES TO

One mis-hire, if it happens
€158,179–€186,179
Assessing that shortlist
€1,490

ESTIMATED GAIN PER HIRE, BEFORE THE FEE

€10,780–€13,176

Smaller than the mis-hire figure, and it should be. That one is a single bad outcome; this is the average across every hire you make: partly a mis-hire avoided, partly a better hire than you would otherwise have made.

AN ESTIMATE FROM YOUR INPUTS AND TWO ASSUMPTIONS OF OURS: AIR ADDS HALF THE PUBLISHED VALIDITY GAP, AND IT IS ONE INPUT AMONG SEVERAL. THE RANGE IS THE SPREAD OF PERFORMANCE FOR THIS ROLE'S COMPLEXITY, INSIDE THE 40-70% OF SALARY THE LITERATURE USES. EVERYTHING RUNS IN ONE CURRENCY: EUR HERE, USD FOR THE US, FEE INCLUDED AT A ROUND CONVERSION. A PLANNING MODEL, NOT AN INVOICE.

WHAT THIS LEAVES OUT

The work that didn't happen while the seat was filled badly. The effect on the people around them. The better candidate you passed over. And the work that was wrong rather than merely late: a Canadian tribunal held Air Canada responsible for what its chatbot told a customer, and US courts have sanctioned lawyers in at least fifteen documented cases since 2023 for filings built on citations a model invented. None of that is in the figure above. Every one of them would make it larger; they are missing because we would have to guess at them.

HUNTER, SCHMIDT & JUDIESCH 1990 · SACKETT ET AL. 2022 · KUNCEL ET AL. 2013 · DELL'ACQUA ET AL. 2026 · BIDWELL 2011 · LACERENZA ET AL. 2017 · PARASURAMAN & MANZEY 2010 · EUROSTAT 2025 · US BLS 2025 where each number comes from

Spread of performance

Hunter, Schmidt & Judiesch (JAP 1990): output varies far more in complex jobs: 19%, 32% and 48% of mean output for low, medium and high complexity. The utility literature denominates the same idea in salary and uses 40% and 70% as its standard estimates. Placing each role on that band is our own interpolation between the two, which is why the answer is a range and not a point.

Hunter, Schmidt & Judiesch 1990 ↗

Employer on-cost

Eurostat 2025 for the EU: non-wage costs are 24.8% of total labour cost on average, 32.3% in France, 23.3% in Germany. US from BLS: benefits are 29.9% of employer cost, so American on-costs are higher than German ones, which surprises most people.

Source: Eurostat, Wages and labour costs (accessed Aug 2026) ↗  ·  US BLS ECEC (accessed Aug 2026) ↗

What better selection buys

Sackett, Zhang, Berry & Lievens (Journal of Applied Psychology 2022), the current authority after they corrected decades of overstated meta-analyses: structured job-specific assessment .42, unstructured interview .19. We use at most half that gap, because AIR is one input to a human decision and not the decision itself.

Sackett, Zhang, Berry & Lievens 2022 ↗

What is counted

The performance shortfall for as long as they stay, the extra management time, and the cost of hiring and onboarding the replacement, the last of those scaled by seniority, from 8% of salary for a support role to 20% for an executive, which is roughly the span SHRM and the CIPD report. Not the original hiring fee or the first ramp: you pay those for a good hire too, so charging them to the mistake would be double-counting. Employer on-costs apply to the time-based items; the spread of performance is denominated in base salary, as the literature defines it.

How much of the role is AI

Dell'Acqua et al. (Organization Science 2026, pre-registered, 758 consultants): on work outside what AI can actually do, people using it were 19 percentage points less likely to be right than people using none. So the deciding variable is judgment about where the tool works, and AIR can only reach the part of a job that depends on it. The full validity gap is discounted twice here: by that share, and by half again because AIR is one input to a human decision rather than the decision.

Dell'Acqua et al., Organization Science 2026 ↗

Why the shortfall persists

Parasuraman & Manzey (Human Factors, 2010) found that automation bias, accepting what an automated aid produces without independent checking, appears in experts as readily as novices, worsens under load, and is not prevented by training or instruction. That is why this model charges the shortfall for every year the person stays rather than assuming they grow out of it. It is a narrow finding about one habit, not a claim that AI skills cannot be taught.

Parasuraman & Manzey 2010 ↗

Picking a team from a workforce

The same Brogden-Cronbach-Gleser arithmetic as the hiring model, with the selection ratio the population gives you: picking twelve from five hundred is sharper than picking one from ten, so it is worth more per head. The horizon is longer because Bidwell (Administrative Science Quarterly 2011) finds internal movers stay where external hires leave, and there is no agency fee, no vacancy and no ramp to charge, because they already work here. Nobody has published a utility analysis of assessment displacing nomination for team staffing: the components are all established, the composite is our application of them. The 18% share net-negative with AI at any one time is our planning estimate; it differs company to company, which is exactly why the model is here to run rather than to quote.

Bidwell 2011 ↗

Aiming development, and why it is the weaker line

The mechanism is evidenced: Lacerenza et al. (2017, 335 samples) find needs-assessed programmes outperform, and the expertise-reversal effect replicates: instruction pitched at the wrong level helps low prior knowledge (d=0.51) and measurably harms high (d=−0.43). The effect size is not. Nobody has priced assess-then-target against blanket rollout. So rather than assume how much more accurately a reading allots development, we derive it: Kuncel et al. (2013) put holistic judgment at r=.28 and mechanical combination of the same evidence at .44, which through Taylor-Russell is an extra 8.2% of the cohort correctly identified. We assumed 0.40 for this once. It was five times too high.

Lacerenza et al. 2017 ↗  ·  Kuncel et al. 2013 ↗

Years they stay

Defaults to four. Average job tenure is about ten years across the EU and 4.8 years for US professionals, so four is deliberately short: a mis-hire who stays as long as everyone else costs considerably more than this model shows.

EVERY INPUT IS AN ASSUMPTION AND EVERY ONE IS ARGUABLE. SAY YOU ALREADY RUN WORK SAMPLES AND THE GAIN ALL BUT DISAPPEARS, WHICH IS THE POINT. A MODEL THAT CANNOT TALK YOU OUT OF IT IS NOT A MODEL.

09 · PLAIN ANSWERS

Asked often, answered straight.

THE ASSESSMENT AT A GLANCE

ASSESSMENTAIR · AI-Readiness Assessment, by VisionVolve
FORMATLive simulation, one sitting; camera and mic on for the recorded close
DURATIONUnder an hour, scored on the spot
LANGUAGESEnglish or Czech
PRICE€149 per person · €110 per person from 100 people · custom quote for larger organisations · excl. VAT
SCORINGFixed, human-authored rules assign every number; models extract what happened, never judge
OUTPUTOne profile per participant, one readout for the organisation
BEST FORTeams and whole organisations that need to know where AI capability actually is
NOT FORStandalone hiring decisions, pass/fail grades, or ranking people

What does AIR measure?

Behaviour with AI, not self-belief: what a person checks, what they trust, and when they stop, captured as roughly forty scored observations across the hour.

How long does a sitting take?

Under an hour, in English or Czech: brief self-ratings, sixteen timed judgment calls, a live incident of about thirty minutes, and a short recorded close.

How much does AIR cost?

€149 per person standalone, €110 per person from 100 people, and custom quotes for larger organisations. Prices exclude VAT.

Is AIR scored by AI?

No score is assigned by a model. Models extract what happened; fixed, human-authored rules assign every number, which is why results land on the spot.

Can AIR be used for hiring decisions?

Not on its own. AIR is not for precise individual grades, pass/fail calls, or standalone hiring decisions; a sitting is one input among several, weighed by someone who knows the person, through a documented human decision.

Is anyone ranked against other people?

Scores compare a person against what their role asks, never against another person. Team-level aggregation is counted, not inferred, and it is never a ranking.