Find your
next edge.
Knowing how to prompt is one part of the picture. Assessor places your AI practice on a five-level ladder, from copying outputs to orchestrating systems, and tells you what the next rung looks like.
The Prompt Crafter
Prompts deliberately and improves results through iteration.
- Uses role, format, and few-shot patterns
- Recognizes common failure modes
- Repairs a failing prompt step by step
More than
prompting.
Seven domains bring knowledge, applied skill, and professional judgment into one picture. Click a point to open a domain. Switch the profile to see how a report shape changes with practice.
Foundational concepts
How language models work, the parameters that change their behavior, and the core architectures behind them.
- Items in the bank
- 20
- Levels covered
- 1–5
- Formats
- 4
Twelve questions.
Each one chosen while you answer.
A fixed test is either too easy for experts or discouraging for beginners. Assessor rates you the way chess rates players. You start at 1000. Answer well and the rating rises, so the next item is harder. Struggle and it falls. Before it chases difficulty, the engine makes sure it has visited all seven domains.
Answer the current item yourself, or set a learner’s true rating and watch a whole session run. Every dot is a real item in the bank, placed by its difficulty.
While any domain is unseen, the next item comes from one of them. Seven of the twelve questions are spent this way.
The remaining five go to the items closest to your rating, where a question tells the engine the most.
Beating a hard item moves you further than beating an easy one. A K-factor of 32 lets twelve answers span the whole ladder.
Competence shows up
in the decisions you make.
Five questions in five formats, one from each of five domains, rising in difficulty. Structured formats are scored by the same rules a full session uses. The open-ended formats, which need a rubric and a scorer, are described below.
Which of the following are valid techniques for improving prompt effectiveness? Select all that apply.
Five items are a taste, not a placement. A full session serves twelve, adapts to you, and includes open-ended tasks.
Recognizing an answer
is not the same as producing one.
Multiple choice measures recognition, the lowest rung of Bloom’s taxonomy. Someone who can define “hallucination” may still miss one in a real output. So seven of the twelve formats ask you to write, critique, design, or reflect, and are scored against a rubric.
Multiple choice
Conceptual knowledge and definitional accuracy.
Scoring. Exact match.
15 in the bankHow open-ended answers are scored
Each open-ended item maps to two to four rubric dimensions, drawn from fifteen shared ones such as accuracy, reasoning, risk awareness, and safeguards. The scorer sees the full one-to-five description of every dimension, reasons before it scores, is told to ignore length, and never computes the final number. Code applies the weights. The prototype calls a local Claude model for this step.
What the report gives you
A level placement with the behaviors that define it, a radar profile across the seven domains, strengths and growth areas, a per-question review with the rubric rationale, learning paths aimed at the weakest domains, and a guide to the behaviors that mark the next level. It is meant to produce a specific next practice, not a number.
What is built, and what still needs validation
The adaptive engine, all twelve renderers, structured scoring, rubric-based AI scoring, session persistence, and the report are built. The bank holds 59 items. Difficulty ratings are expert estimates awaiting pilot calibration. Human–AI scoring reliability and a fairness analysis across linguistic backgrounds are planned. The framework is for self-directed development, not hiring, performance review, or credentialing.