A perspective, in brief.
Why static measures of AI proficiency are becoming inadequate for increasingly hybrid professional roles
The rapid diffusion of generative AI into professional work has created a curious problem.
Organizations increasingly want employees to become “AI-ready.” Professionals describe themselves as AI-fluent, AI-native or AI-first. Job descriptions now routinely ask for familiarity with LLMs, automation, AI-enabled workflows or applied AI.
Yet the underlying construct remains surprisingly imprecise.
What, exactly, does it mean for a professional to be AI-ready?
The question appears simple, but becomes difficult as soon as we move beyond specialist AI roles.
An ML Engineer, a Product Manager, a Transformation Leader, a Designer and a Business Head may all require meaningful AI capability. But the capability expected of each is structurally different.
The situation becomes even more complex because modern professional roles increasingly overlap.
A Product Manager may be expected to prototype using APIs. An Engineer may need to understand commercial priorities and user behaviour. A Transformation Leader may need sufficient technical depth to distinguish a viable AI workflow from a superficial automation proposal. A senior executive may need to understand model risk, governance, evaluation and organizational adoption without ever writing production code.
The conventional question—
“How capable is this person at AI?”
—may therefore be less useful than another:
“How well does this person’s AI capability align with what this particular role requires?”
That distinction is the premise behind a framework I have been developing: ROSAI — Role-Oriented Skills Assessment for AI.
AI capability and AI readiness are not the same construct
The central proposition of ROSAI is deliberately simple:
AI capability belongs to the individual. AI readiness exists in relation to a role.
An individual may possess a certain capability profile across areas such as AI fluency, applied judgement, technical execution, evaluation, governance and business impact.
That profile belongs to the individual.
But whether that profile represents “strong AI readiness” depends on what the individual is expected to do.
Consider someone with strong AI judgement, high business orientation, good evaluation discipline and moderate technical execution, but limited depth in traditional machine learning.
That profile may be highly appropriate for an AI Product Manager or Transformation Leader.
It may be substantially less appropriate for a role centred on production ML engineering.
Nothing about the individual changed.
The interpretation changed because the role changed.
This distinction matters because most professional assessments implicitly assume that the thing being measured has a relatively stable meaning across contexts.
For AI readiness, that assumption may be increasingly difficult to defend.
The limitations of job titles as assessment categories
One obvious response is to create separate assessments for different professions.
Product Managers receive one assessment. Engineers another. Data professionals another.
That helps, but still leaves a problem.
Job titles are becoming weak proxies for actual responsibility.
Two professionals with the same title may operate in fundamentally different environments.
A Product Manager in one organization may primarily own discovery, prioritisation and commercial outcomes.
Another may work deeply with APIs, prompt orchestration, retrieval systems and model evaluation.
A third may be responsible for AI transformation across several business functions.
The title remains the same.
The work does not.
ROSAI therefore does not begin by assigning an individual to a single professional category. Instead, the target role is represented as a continuous composition of responsibilities across six broad families:
Product / Business, Software Engineering, Data / ML, Operations / Transformation, Design / UX, and Leadership / Governance.
A role might therefore be represented as 50% Product / Business, 20% Software Engineering, 15% Transformation, 10% Leadership and 5% Data / ML.
Another role with the same job title could have a very different composition.
The purpose is not to claim mathematical precision about how a job is constituted. The purpose is to preserve something that categorical job titles often lose:
the fact that modern roles are mixtures.
From role composition to capability importance
Once the role is represented as a responsibility mix, the next question becomes:
Which AI capabilities matter most for that mix?
ROSAI currently models professional AI capability across seven dimensions:
AI Fluency captures practical understanding of AI concepts, capabilities, limitations and common solution patterns. Applied AI Judgement concerns decisions about when and where AI should be used, including trade-offs involving accuracy, cost, latency, privacy, risk and human oversight. Build & Execution concerns the ability to move from an idea toward a prototype, workflow, integration or production implementation. Evaluation & Reliability concerns whether an AI system can be assessed systematically rather than accepted because its outputs appear plausible. Data & ML Capability represents the level of data and machine-learning depth required by the role. Responsible AI & Governance covers issues such as privacy, permissions, security, fairness, explainability, auditability and oversight. Business & Organizational Impact concerns the ability to connect AI activity to adoption, productivity, customer outcomes, workflow redesign, cost, revenue and organizational change.
The role composition then determines the relative importance assigned to these capabilities.
At a high level:
Role composition → Capability weights → Role-relative interpretation
This is where the framework moves away from a static proficiency model.
The same capability profile can produce different readiness interpretations under different role compositions because different capabilities matter to different degrees.
Why a single score remains insufficient
Even a role-adjusted score creates another problem.
Suppose two people receive the same readiness score.
One has built live systems, documented architecture, produced measurable business outcomes and can provide credible evidence of prior work.
The other has little supporting evidence.
Should the two profiles be interpreted identically?
Probably not.
But incorporating evidence directly into the capability score creates a conceptual problem of its own.
Evidence of capability and capability itself are related, but they are not identical constructs.
A professional may possess substantial capability but be unable to disclose confidential employer work.
Another may have a highly visible public portfolio but weaker depth than the visibility implies.
ROSAI therefore separates three signals.
ROSAI Score
The role-adjusted estimate of AI capability.
It addresses:
How well does the assessed capability profile align with the requirements of this role?
Evidence Strength
The strength of supporting evidence behind the capability claims.
This may include live applications, repositories, demos, architecture documents, published work, employer projects, measurable outcomes or references.
Assessment Confidence
The degree of confidence with which the available assessment signal can be interpreted, considering factors such as completion, internal consistency, adaptive scenario responses, response quality and supporting evidence.
The distinction is deliberate.
Evidence Strength and Assessment Confidence do not directly increase or decrease the ROSAI Score in Version 1.0.
They provide context around it.
This avoids collapsing capability, proof and interpretive confidence into one number.
The problem of compensation
Weighted scoring introduces another methodological issue.
Strong performance in one area can mathematically compensate for weak performance elsewhere.
In many contexts that is reasonable.
In others, it is not.
Consider a role in which evaluation and reliability are central. A professional may be highly fluent in AI terminology, excellent at identifying use cases and strong at business prioritisation. But if that person has very weak understanding of evaluation, validation and failure handling, should excellence elsewhere completely neutralise that weakness?
For some roles, that would produce a misleading interpretation.
ROSAI therefore incorporates the principle of role-critical safeguards .
The principle is:
A severe gap in a capability that the role itself identifies as core should not always be averaged away by unrelated strengths.
The public framework describes this principle while leaving the exact operational thresholds and answer-level coefficients undisclosed.
That separation is intentional. Methodological transparency does not necessarily require publishing an assessment’s complete answer key.
Assessment should test judgement, not merely familiarity with terminology
Another design question concerns what an AI-readiness assessment should actually ask.
Self-reported questions such as “How advanced are you at AI?” are easy to administer but weak as primary signals.
Terminology-based tests also create problems. AI terminology evolves rapidly, and recognition of vocabulary does not necessarily demonstrate applied capability.
The current SCORE-AI implementation of ROSAI therefore combines several kinds of signal.
Structured questions examine areas such as AI usage, decision factors, validation approaches, production readiness, practical implementation experience and Data/ML maturity.
These are complemented by short adaptive scenarios based on the respondent’s role composition.
The scenarios are evaluated against four broad criteria:
Context Framing, Decision Quality, Evaluation and Control Discipline, and Role-Specific Execution Depth.
The intention is not to reward sophisticated vocabulary.
A concise answer demonstrating clear reasoning should outperform a jargon-heavy answer that fails to address the underlying problem.
This becomes particularly important in an era when respondents themselves can use AI to formulate polished answers.
Assessment design therefore needs to become less dependent on verbal polish and more dependent on decision structure, consistency and applied reasoning .
Formalisation is not validation
This distinction deserves particular emphasis.
A framework can contain equations, scoring models, defined constructs and structured rubrics and still remain empirically unvalidated.
Mathematical notation makes assumptions explicit.
It does not make those assumptions correct.
ROSAI Version 1.0 should therefore be understood as a practitioner-designed framework and testable methodology , not as a psychometrically validated employment instrument.
The value of formalising the framework is different.
It creates something that can be examined.
Weights can be challenged.
Constructs can be tested.
Scenario grading can be measured for agreement.
Role sensitivity can be evaluated.
Bias can be investigated.
The framework can therefore move from intuition toward evidence.
The proposed validation path includes expert content review, pilot data collection, reliability testing, grader-agreement analysis, construct validation, known-groups analysis, criterion-related validation, fairness analysis and eventual recalibration.
Why publish the methodology rather than only the product?
ROSAI emerged while I was building SCORE-AI , a working assessment product based on the framework.
It would have been considerably easier to publish the application, describe it as “AI-powered” and leave the scoring architecture opaque.
But that creates an increasingly common problem in AI products:
the interface becomes visible while the reasoning model remains invisible.
If an assessment produces a number that may influence how professionals understand their capabilities, the underlying assumptions should at least be available for scrutiny.
For that reason, I documented ROSAI separately from SCORE-AI.
The framework paper sets out the conceptual model, role-relative weighting architecture, capability dimensions, interpretation model, role-critical safeguards, limitations and proposed validation process.
I have made that methodology publicly available so that the assumptions behind the assessment can be examined independently of the product itself.
ROSAI Version 1.0 is now available through public research records of the same framework paper; their availability should not be interpreted as peer review or empirical validation.
ROSAI Framework https://score-ai.prasanjitsaha.com/rosai-framework
Framework paper —
https://doi.org/10.5281/zenodo.22801326
https://doi.org/10.2139/ssrn.7479102
Working implementation — SCORE-AI https://score-ai.prasanjitsaha.com/
SCORE-AI therefore functions as the first working implementation of ROSAI rather than being the framework itself.
The larger issue is organizational, not merely individual
The question of professional AI readiness extends beyond hiring.
Organizations are now trying to determine how to develop AI capability across entire workforces.
Universities are reconsidering what graduates need to know.
L&D functions are designing AI training programmes.
Leaders are deciding which capabilities should remain specialist and which should become broadly distributed.
Professionals are trying to understand what they personally need to learn.
A generic instruction to “become better at AI” is unlikely to be sufficient.
The more useful question is:
Which AI capabilities matter, to what depth, for which responsibilities?
That shifts the conversation from AI literacy as a generic attribute toward AI capability as an organizational design problem .
It also implies that AI upskilling should perhaps be designed around responsibility architecture , not just around job titles or standardized course catalogues.
A finance leader, product leader and operations leader may all need substantial AI capability.
But they may need very different capability profiles.
That has implications for hiring, internal mobility, learning design, workforce planning and leadership development.
And that, to me, is the more interesting question behind ROSAI.
A framework should invite disagreement
ROSAI Version 1.0 is intentionally not presented as a finished answer.
Several assumptions remain open to challenge.
Are the seven capability dimensions the right ones?
Are some dimensions too broad?
Should evidence remain completely separate from capability?
How should role-critical weaknesses be treated?
How stable are role-composition weights across industries?
Can short scenario responses provide a useful applied signal?
Do different evaluators interpret the same scenario consistently?
These are empirical questions.
They should eventually be answered with data rather than assertion.
My objective in publishing the framework is therefore not simply to propose another AI-readiness score.
It is to make the underlying assumptions explicit enough that they can be discussed, tested and improved.
Because as AI becomes embedded into more professional roles, the question is no longer merely whether people know how to use AI.
The more consequential question may be:
Do they possess the right combination of AI capabilities for the responsibilities they are expected to perform?
That is the question ROSAI is designed to explore.
Explore the ROSAI Framework https://score-ai.prasanjitsaha.com/rosai-framework
Try the SCORE-AI implementation https://score-ai.prasanjitsaha.com/
About the author
Prasanjit Saha works across Digital Product Management, Enterprise Transformation and Applied AI, with a focus on converting emerging technologies into practical products, workflows and measurable business outcomes.
He developed ROSAI — Role-Oriented Skills Assessment for AI and built SCORE-AI as its first working implementation.
Key themes
- AI capability belongs to an individual; readiness exists in relation to a role.
- Modern roles are mixtures, so job titles are weak proxies for assessment.
- Capability, supporting evidence and confidence should not be collapsed into a single number.
- Formalisation makes a framework testable; it does not make it validated.