Skip to main content
Line art on a black background: a completed checklist beside a magnifying glass, a question mark and a circle of measuring points.

AI & Automation

Testing AI Skills: How Well Can You Really Use AI?

You use ChatGPT every day — but how well do you actually understand AI? An AI skills test is only worth taking if it measures more than prompting: tool choice, data protection, failure modes, building your own agents and cost all belong in it.

Michael Schranz · AHEAD OF TIME13 min

An AI skills test should measure more than whether someone has heard of ChatGPT or can write a usable prompt. What matters for SMEs is whether people apply AI to real work situations appropriately, critically, safely and economically. That includes fundamentals and AI literacy as much as prompting, tool choice, data protection and source checking — and at the upper levels agents, automation, governance and the question of what all of it costs in tokens and money.

But an assessment is only useful if it defines clearly what it wants to measure, spreads its questions across different competence levels and reduces typical test biases. For AI evaluations, NIST stresses that measurement is always context-dependent and needs valid tasks, metrics and datasets. On AI literacy, the European Commission likewise avoids a one-size-fits-all approach: prior technical knowledge, experience, role and usage context all belong in the picture (European Commission, 2026; NIST, 2026a).

The AI skills check from AHEAD OF TIME is exactly such a test — built across six independent levels, from first steps to the strategy layer. The result is not meant to be a label but an orientation: what can you already do, and what is worth learning next?

What this is about

"I work with ChatGPT every day." That sounds like solid AI competence. Maybe it is. Maybe it only means someone regularly has texts rephrased.

Frequency of use and competence are not the same thing. A person can work with generative AI daily and still struggle to verify sources, handle confidential data correctly, tell RAG from fine-tuning, or pick the right tool for a task. Conversely, someone can explain technical concepts cleanly and still barely use AI productively in day-to-day work.

A sensible AI skills test makes those differences visible. That is where its value for SMEs lies: not ranking people against each other, but recognising competence profiles and learning needs.

Why daily AI use is no proof of AI competence

Generative AI has lowered the entry barrier radically. A natural-language interface quickly creates the feeling of having understood the system. But good usability hides real complexity.

Professional use raises questions like: which information am I even allowed to enter? When do I need a source? When is a model with web access better than a source-grounded system? What happens when an agent is allowed to call external tools? Is a convincingly phrased answer also correct? And do I really need the largest reasoning model for this task?

For AI literacy, the European Commission names exactly these context-related considerations: organisations should understand which AI systems they deploy, which opportunities and risks come with them, and what knowledge employees need given their role and experience (European Commission, 2026).

Go deeper: Part 1 of this series — AI skills in SMEs explains which capabilities leaders and teams actually need in 2026.

What an AI skills test should measure: nine topics

An assessment should not try to capture "AI competence" with a single knowledge question. A multi-dimensional model works better. The AI skills check examines nine topics — and the same nine at every level. What changes is not the set of topics but the depth of the questions.

Can someone assess the opportunities and limits of generative models realistically? Do they understand why a model hallucinates, why a linguistically confident output is not automatically true, and what distinguishes an agent from a chat window? This is the base layer — without it, every later decision is a lucky guess.

2. AI tool landscape

AI competence shows in picking the right tool for a specific use case. Source-grounded analysis, current web research, coding, image generation, voice or agentic automation all call for different systems. What gets assessed is the fit to the described case — not whether a brand currently counts as "the best".

3. Data protection & sovereignty

Does someone recognise when sensitive data is involved? Is it clear where data is processed, who can access it and what that means for a Swiss SME? With more advanced systems, permissions, prompt injection and human-in-the-loop come on top. NIST places security, robustness, transparency and risk management among the core dimensions of trustworthy AI (Autio et al., 2024).

4. Prompt & context engineering

Can a task be described so that goal, context, requirements and format are clear? More importantly: does the person recognise when a better prompt is enough — and when the real problem is missing data, an unsuitable tool or a broken process?

Practical knowledge on this: prompt engineering for SMEs.

5. Tool-specific knowledge

Tool choice is one half, detailed knowledge the other. What can a model do with which context window, where are the limits of a file upload, where does web search end, what does a project or custom-GPT feature actually do? This knowledge ages fast, which is why it needs regular review — and why it is a topic of its own rather than being folded into the tool landscape.

6. Failure modes & limits

Hallucination is only one kind of error. Add stale training data, truncated context, silently dropped intermediate steps, convincing false citations, and automations that keep running even when something has gone wrong. Whoever knows the typical breaking points checks in the right place instead of checking everything equally.

7. Building your own agents

Advanced users know more than chat interfaces: they understand how AI systems are connected to knowledge, tools and external services — RAG, tool calling, MCP, automation chains. What counts is less the ability to program every technology yourself than the ability to judge its possibilities, limits and risks.

More practical examples: the AI & automation topic cluster.

8. Vibe coding & setup

AI writes code — and a growing share of the work in SMEs is created that way: small tools, prototypes, scripts, automations, built by people without a developer role. That is a competence field of its own. It includes getting a setup running, reviewing generated code instead of adopting it blindly, and knowing when a prototype has to stay a prototype.

9. Efficiency & cost

Professional use also means weighing quality against resources. Anyone who uses unnecessarily long contexts, oversized models or superfluous agent steps in token-based APIs produces cost and compute without any gain in quality.

For companies: AI transformation, education & engineering at AHEAD OF TIME.

Six levels instead of "pass or fail"

A binary result — competent or not competent — does not help anyone learn. The AI skills check therefore structures competence across six independent levels. Each has its own question catalogue, and each tests the same nine topics at a different altitude.

LevelNameWhere that puts you
1AI BeginnerYou know individual terms and tools, but the overall picture is still forming. The biggest lever right now is the fundamentals: what a language model actually does, where it is reliable and where it is not.
2AI PractitionerYou already work with AI tools and make workable day-to-day decisions. What is still missing is the systematic side: which tool when, and how to verify results reliably.
3AI AdvancedYou have an overview of the tool landscape, know the typical failure modes and can judge results properly. The next step is towards building: your own workflows, context and guardrails instead of one-off requests.
4AI IntegratorYou no longer use AI tool by tool but connect tools into workflows: data moves from one system to the next and steps build on each other. The next lever is making those chains reliable and traceable.
5AI ArchitectYou build things yourself — agents, automations, data paths. You know not just the tools but where they break: permissions, retries, failure modes. The next step leads away from technology towards the organisation.
6AI StrategistYou think about AI from the organisation outwards: who is allowed to do what, what a decision rests on, which dependencies a company takes on. This is the level at which AI either carries its weight or becomes expensive.

The difference to a single overall score matters: an average across all topics hides strengths and weaknesses. Someone who shines on the tool landscape and fails on data protection lands on "medium" overall — and learns nothing from it.

How a run works

Each level holds 24 to 26 questions spread across the nine topics. One level takes 10 to 15 minutes. Anyone reaching 80 percent has passed and unlocks the next one — level 6 has no successor, so there the threshold simply means: done.

Three things are deliberately built this way:

  • Every answer is explained. Not just "right" or "wrong" but why — including why the obvious wrong option looks so obvious. That makes the test a learning format in itself.
  • The evaluation is per topic. You do not only see a percentage but which of the nine fields is the weak one.
  • Every level comes with a learning set as a PDF, containing exactly the topic modules tested at that stage — meant as preparation, not as homework.

Answer options are shuffled on every page load. The letter position carries no meaning, and memorising an order teaches you nothing.

Start the AI skills check now → · Overview of both checks

Why hard questions and good distractors matter

A test is not automatically good just because it feels hard. Difficulty also comes from unclear wording, niche knowledge or trick questions. In a competence test, items should instead reflect relevant decisions and separate different levels of knowledge and ability.

In multiple-choice questions, the wrong answer options — the distractors — do the decisive work. They have to be plausible and drawn from related concepts. Among other things, the American Psychological Association recommends deriving distractors from relevant subject content, avoiding conspicuous length differences for the correct answer, and dropping constructions such as "all of the above" or unnecessarily negative phrasing (American Psychological Association, 2018).

The chance component belongs in the calculation too. With four answer options, blind guessing hits 25 percent of the time. That is why no single hard question ever decides a classification. What carries meaning is the pattern across 24 to 26 items and nine topics.

How test bias can be reduced

In practice, no assessment is entirely bias-free. The more honest formulation is therefore: reduce bias systematically and verify it empirically. Five points are central.

  • Answer-position bias: the correct answer must not sit conspicuously often in the same position. In the web app the order is randomised, and the correct answer is stored via a stable ID instead of A/B/C/D.
  • Length and detail bias: the correct answer must not be regularly longer, more precise or more fluently technical than the distractors.
  • Distractor bias: obviously absurd options turn a knowledge question into an elimination exercise. Wrong options have to stay technically plausible.
  • Vendor and familiarity bias: on tool questions, brand recognition must not be mistaken for competence. Scenario and selection criteria weigh more than the product name.
  • Construct-irrelevant bias: a question should measure the intended AI knowledge — not reading speed, English proficiency, arithmetic tricks or experience with one particular user interface.

Once enough answers exist, the real quality check begins. Item difficulty, discrimination, distractor efficiency and response times show which questions actually work. With a large enough sample, differences between roles, age groups or language versions can be examined too. NIST stresses that reliable AI evaluation should explicitly account for measurement goals, context, uncertainty and item difficulty (Keller et al., 2026; NIST, 2026a).

What a test result can do — and what it cannot

An online test is an orientation, not a certification of professional competence. Multiple-choice items assess knowledge, understanding and decision-making efficiently. They do not show how well someone sets up a complex AI workflow in a real work context.

For SMEs, the result is most valuable when it serves three functions:

  1. Orientation: where do the person or the team stand today?
  2. Diagnosis: in which of the nine topics are the concrete gaps?
  3. Next learning step: which content, exercises or training make sense next?

NIST follows the same basic idea in AI evaluation: measurement should make capabilities and limits visible and always be interpreted in the respective application context (NIST, 2026a).

From result to learning path

A score alone changes no way of working. The value appears when assessment and learning are connected:

Assessment → skill gap → learning content → practice → retest

Someone who does well on the tool landscape but shows gaps in data protection or in building their own agents does not need a generic "AI for beginners" course. And someone at level 1 should not be run over with MCP specifications or multi-agent architectures.

That is exactly where a competence test becomes strategically interesting for SMEs: it enables differentiated learning instead of training by watering can. AHEAD OF TIME connects this logic with custom AI training, team masterclasses and AI transformation — from skill check to application on real company tasks.

Part 3 of this series shows how a test result becomes a learning path for management, business functions and technical roles.

Conclusion: the score is not the point

A good AI skills test answers more than "how many questions were right?". It shows which capabilities exist and where the next learning investment pays off most.

For SMEs that means: fewer blanket AI courses, less overconfidence from mere tool usage, a better basis for roles, governance and training. At the same time the test itself has to stay capable of learning. Tool questions age, new competence fields appear, items need reviewing against real usage data.

The check is therefore not the end of competence development but its starting point.

Q&A — testing AI skills

What is an AI skills test?

A structured assessment that examines knowledge and applied capabilities around artificial intelligence. A good test covers several topics — from LLM fundamentals, prompting and tool choice through data protection and failure modes to building your own agents, vibe coding and cost.

How many questions does the AI skills check have?

24 to 26 questions per level, spread across nine topics. One level takes 10 to 15 minutes. Anyone reaching 80 percent unlocks the next of the six levels.

Which levels are there?

Six: AI Beginner, AI Practitioner, AI Advanced, AI Integrator, AI Architect and AI Strategist. Each level is its own question catalogue and tests the same nine topics in greater depth.

Can an online test measure real AI competence?

It measures knowledge, understanding and decision-making ability — it does not replace observing actual work performance. Its value lies in the orientation it gives and in the learning path that follows.

Why is testing prompting alone not enough?

Because productive AI use additionally requires tool selection, quality control, data protection, source literacy and, depending on the role, technical concepts such as RAG, agents or automation. Prompting is one of nine topics, not the whole field.

What does the top level mean?

AI Strategist does not mean knowing every product. It means being able to think about AI from the organisation outwards: permissions, decision bases, dependencies, risks and evaluation.

How do you prevent bias in multiple-choice questions?

Through plausible distractors, randomised answer positions, comparable answer lengths, clear language and avoiding brand recognition as a hidden competence indicator. None of that makes a test entirely bias-free — it makes it verifiable. After launch, items are additionally analysed empirically.

How often should an AI skills test be updated?

Fundamental concepts and security principles change slowly. Tool-specific knowledge and product questions need reviewing far more often, because features and market positions shift quickly.

Does an AI skills test make sense for whole teams?

Yes — as long as the results are not used as a ranking. Aggregated competence profiles show where shared fundamentals are missing and which roles need targeted depth.

Link copied

Let's talk — no sales pressure.

Tell us about your project. We'll come back with a concrete answer, not a sales email.

Send an enquiry

A form, two minutes. We reply within one working day — or book a slot directly.

Rather book a 30-min call
Link copied