Entrance
Occasion
What is coming up?
A real appointment, in one sentence. That turns a good intention into preparation with a deadline — and a vitamin into a painkiller.
Methodology · AI coaching
An AI role-play is a conversation. AI coaching is a system that decides which conversation comes next. The Careertrainer Loop is the model behind it: five stations, a foundation, a yield. This page explains how it works, how the evaluation inside it operates, and where the limits are.
The loop at a glance. Each station is described on its own below.
01
Three properties that turn separate conversations into coaching.
A catalog provides scenarios and leaves the choice to chance or preference. A coach decides what comes next — and explains why.
Not only within one conversation, but across many — and across different counterparts of the same type. Only then can a pattern be recognized at all.
A catalog asks what you want to practice. The loop says what is due and why — based on what became visible in recent conversations.
Not “practice is good”, but: “Thursday proposal meeting, procurement is blocking on price.” That turns an intention into preparation with a deadline.
02
Five stations, a foundation, a yield.
Entrance
What is coming up?
A real appointment, in one sentence. That turns a good intention into preparation with a deadline — and a vitamin into a painkiller.
Station 1
How does it actually go?
The live conversation. Two separate AI systems: one plays, one scores. The character does not know the scoring standard and therefore cannot be gamed. During the conversation only one sentence is on screen: the goal.
Station 2
What was it down to?
Self-assessment before the result, then the evaluation with quotes from your own transcript — and the gap between the two. Disagreement is intended: an evaluation you cannot argue with is a verdict.
How it works in detailStation 3
What is different this time?
Not a second attempt at the same scenario, but a named variant: tighter, harder, or transferred to another person. The change is fixed in advance.
Exit
What do you take with you?
One sentence before the real appointment — and afterwards the question of whether the assumption held. That closes the circle back to the occasion.
Foundation
What makes the conversation yours?
Company context, product or guideline, persona. Without these three layers every scenario stays generic. The part that grows more valuable over time: after months, objections from real conversations sit there.
Yield
What shows up over time?
Not a score, but a statement about recurring behavior toward a certain type of counterpart. Aggregated at team level, that becomes the answer to falling back into the old pattern.
What a pattern sounds like
“With quiet market leaders you give in on the third follow-up. With time-pressure types you do not.”
An illustrative wording, not an evaluation of a real person. Aggregated at team level, the same statement might read: “Your team wins against the price-driven buyer and loses against the waiting decision-maker.”
03
Whoever plays the role does not score it.
During the conversation, one model takes the role of your counterpart. Afterwards, a second, independent model reads the transcript and scores it. The model that played the role does not evaluate the conversation itself.
The evaluating model is not allowed to assign an overall grade. It scores each goal and each competency separately. The system then calculates the overall grade, using a weighting that is the same for every conversation.
The model does not invent its own standards. What should be achieved in a scenario, and how that shows up, is set when the scenario is created — by people, before the conversation. After that it stays fixed.
So far each conversation is scored on its own. With the AI coach, the history is added: the evaluation sees what the same person showed in earlier conversations and can name a pattern instead of only a snapshot. That is the prerequisite for the stations Repetition and Transfer — and the part of the loop that is still being rolled out at the time of this page.
It does not change the scoring of a single conversation: weighting, levels, and limits stay as described. Earlier conversations feed the recommendation, not the grade.
Two small rules belong with this: the first sentence comes from the model, not from you, so it is not scored. And passages where speech recognition has obviously misunderstood something are left out.
70%
Goals of this scenario
30%
Competencies
Set for each scenario: the situation, what counts as achieved, and how heavily it weighs.
A fixed set for each training area.
Every individual score uses the same scale
8–10
shown reliably
6–7
visible, not consistent
4–5
partial
0–3
barely shown
The competencies depend on the training area — what matters in a negotiation is different from what matters in a leadership conversation.
04
A level that only says “good” means something different to everyone.
When the behavior is described concretely, everyone shares the same reference point. The room for interpretation gets smaller.
That is why each competency is described across four written levels. We do not score who someone is. We score what was visible in this situation. Two examples from different training areas:
Example 1 · Leadership · “active listening”
8–10 Asks targeted questions and restates what came across.
Picks up earlier points later. The other person feels understood.
6–7 Listens and asks follow-ups, but not throughout.
Some signals from the other person go unaddressed.
4–5 Asks questions but barely hears the answer.
Their own agenda drives the conversation.
0–3 Talks alone for long stretches or interrupts.
No dialogue forms.
Example 2 · Sales · “handling objections”
8–10 Acknowledges the objection and asks what is behind it.
Then addresses that point directly. The conversation continues.
6–7 Responds to the objection, but stays on the surface.
The underlying reason stays unclear.
4–5 Justifies themselves or repeats their own argument.
The objection is still sitting there afterwards.
0–3 Skips the objection or gives in immediately.
The substance is never discussed.
Illustrative only. Each training area has its own competencies and its own levels — and all of them can be adapted to your organization.
Every judgment is tied to a passage from the conversation. That makes it possible to see what it refers to — and to disagree, if needed.
Active listening
6 of 10
Passage from your conversation
“I understand. Let’s still stick with Friday — I need the numbers by then.”
Two sentences earlier your counterpart had said that several things were happening at once. You heard it and acknowledged it, then went straight back to your deadline. A short follow-up question would have been enough to learn what it actually depended on — and whether Friday was realistic.
An example, reconstructed from a leadership scenario. Quotes always come from the person’s own conversation; invented evidence is not allowed for the evaluating model.
05
We made each of them on purpose. None is a side effect of the technology.
Every piece of feedback is tied to a specific moment in your conversation and describes what happened there — not who you are.
Feedback works while attention stays on the task. When it shifts to the person, the effect fades. In a third of the cases studied, feedback even reduced performance.
Kluger & DeNisi, 1996
Four written levels per competency instead of abstract grades. Each one states what was visible.
Scales defined by observable behavior give every rater the same reference point.
Smith & Kendall, 1963
Every scenario has its own goals: a concrete situation in which it is clear what counts as achieved. Each goal is weighted and scored on its own.
Scoring is based on concrete situations from everyday work, not on general traits.
Flanagan, 1954
The evaluating model assigns individual scores, but not an overall grade. The weighting is fixed in the system.
It is always possible to see how a grade was produced — and it is produced the same way across conversations, groups, and time.
Design decision
Four constraints keep a rating from tipping — neither into harshness nor into vagueness.
3 points
The maximum deduction for unfavorable conversation patterns — such as giving in too early, getting personal instead of staying on the issue, or leaving without a firm outcome. A conversation never loses more than that.
3 contributions
That many intelligible contributions are required at minimum. Below that there is no grade at all: a rating on a thin basis is worse than none.
70/30
Scenario goals count for 70 percent, competencies for 30. The model cannot change that — which is why evaluations can be compared with one another.
2 models
One talks with you, the other evaluates. Whoever played the role does not judge the scene.
Feedback that, after one mistake, only shows the failure no longer tells anyone what to do differently next time.
06
The foundation of the loop is the part you fill.
Do you work with your own conversation guide, leadership principles, or competency model? After we agree on it, we place it into the evaluation. Participants then receive feedback in your language and against your standards, rather than generic ones.
What you bring
Your leadership principles
or competency model
Your conversation guide
if you use one
Real situations
from the everyday work of your teams
What we set up with it
Competencies use your terms
Written the way people in your organization talk about conversations.
Each scenario gets its own goals
Drawn from situations that actually occur for you.
Your model becomes the standard
Scoring follows your framework, not someone else’s.
Tone and boundaries fit you
You supply the material and the terms. We do the setup.
If you bring nothing of your own, that is fine too. The scenarios work out of the box without a particular model.
Negotiation training is an exception. Terms such as best alternative, anchoring, and concession are fixed parts of the competencies — without them a negotiation cannot be scored in a meaningful way.
The 70 to 30 weighting cannot be changed. That is intentional: only then do evaluations stay comparable across teams and across months.
07
Four things an evaluation cannot do.
We name them because they matter more for interpretation than any strength we could list.
Not a suitability assessment.
The rating is not a test and is not a basis for personnel decisions. It is a training instrument.
No validated comparison with human raters.
Our scale is built on the principle of behaviorally anchored rating. We have not yet measured how closely it matches the judgment of experienced trainers.
No individual results for managers.
The result belongs to the person who practiced. Reporting upward happens only at group level — including the patterns from the loop.
No result that repeats word for word.
The evaluation runs with very little randomness and in a fixed format. The same conversation scored twice yields the same judgment, but not the same text word for word.
You can test the second point in live use: two experienced people score a sample of transcripts against the same levels, without seeing the system’s scores. The comparison shows where the two judgments diverge. The effort is about a day, and the result is a defensible number instead of an assumption. Talk to us if you want to plan that for your program.
Smith & Kendall (1963)
Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales.
Journal of Applied Psychology, 47(2), 149–155.
Foundational work on behaviorally anchored rating scales. Evidence that they reduce rating bias is positive, but not consistent.
Kluger & DeNisi (1996)
The effects of feedback interventions on performance.
Psychological Bulletin, 119(2), 254–284.
Meta-analysis of 607 effect sizes. Feedback improves performance on average, but reduces it in more than a third of the cases studied.
Flanagan (1954)
The critical incident technique.
Psychological Bulletin, 51(4), 327–358.
A method for deriving rating criteria from concrete professional situations.
Questions about the methodology? We are glad to answer them in detail — including for works councils, data protection, or specialist teams.
Careertrainer.ai[email protected]As of September 2026