1
Occasion
Start from a concrete conversation
The starting point is a real, upcoming, or role-typical conversation. People then practice situations in which they want to improve how they lead the conversation, rather than a scene detached from work.
Methodology · AI coaching
Conversation simulations create practice. Training depends on what happens before and after: which situation is practiced, how the conversation is reflected on, what should be trained next, and how that reaches everyday work. The Careertrainer Loop connects these steps into a training process.
The training process at a glance. Each step is explained below.
01
A scenario library mainly answers what someone wants to practice today.
The Careertrainer Loop goes further. Earlier conversations and recognized development points can be used to suggest a fitting next exercise. Separate simulations then become training that builds on itself.
Exercises are not treated as fully isolated. Across several conversations it can become visible where strengths and difficulties repeat, including with similar roles.
A library leaves the choice to participants. A coach can also take into account which situations have already been trained, where it got difficult, and which exercise fits that.
The starting point is a real conversation, or one that is typical for the role — for example Thursday’s proposal meeting, where procurement is blocking on price. The exercise then has a link to everyday work.
02
Five steps, company context as the foundation, and patterns as the result of several runs.
1
Start from a concrete conversation
The starting point is a real, upcoming, or role-typical conversation. People then practice situations in which they want to improve how they lead the conversation, rather than a scene detached from work.
2
Play the conversation in a realistic setting
The participant holds the conversation with an AI persona. Its behavior is tuned to the scenario, the role, and the situation, and it reacts to how the conversation develops. Simulation and evaluation are separate: the persona holds the conversation, and the evaluation follows afterwards from the full transcript.
3
Bring self-assessment and evaluation together
Before the evaluation appears, the participant assesses the conversation. Feedback on conversational behavior follows, with concrete passages from the transcript. Self-assessment and evaluation can then be compared. The evaluation is a reasoned reading against defined criteria, not a final verdict.
How it works in detail4
Practice again under changed conditions
The next exercise does not simply repeat the same scenario. Individual conditions change, for example the other person’s reaction, how sharp an objection is, or the personality type. In the next simulation the customer may push harder on price, or an employee may dodge the feedback more. Earlier exercises can be used to suggest what should be trained next.
5
Apply what was learned at work
Training does not end with the simulation. Before a real conversation, people can note what they want to apply. Afterwards they can check whether the approach worked and what that means for further exercises.
Foundation
Use the actual work context
A useful simulation needs more than a generic role description. Company, role, and the concrete situation can be taken into account: products, internal guidelines, typical objections, conversation goals, or relevant counterparts. The exercises then sit closer to the conversations people actually have in the company.
Pattern
See development beyond a single conversation
One role-play is a snapshot. Across several exercises, recurring strengths and difficulties can become visible — in certain phases, with objections, or with certain kinds of counterparts. These patterns can be used to choose further training steps more deliberately.
What a pattern sounds like
“With quiet market leaders you give in on the third follow-up. With time-pressure types you do not.”
An illustrative wording, not an evaluation of a real person. Aggregated at team level, the same statement might read: “Your team wins against the price-driven buyer and loses against the waiting decision-maker.”
03
Simulation and evaluation are separate.
During the simulation, one model takes the role of the counterpart. Afterwards, a second, independent model reads the transcript and evaluates it. The persona that held the conversation does not evaluate it. The role stays in the conversation, and the evaluation refers to the full transcript.
The evaluating model is not allowed to assign an overall grade. It scores each goal and each competency separately. The system then calculates the overall grade, using a weighting that is the same for every conversation.
The model does not invent its own standards. What should be achieved in a scenario, and how that shows up, is set when the scenario is created — by people, before the conversation. After that it stays fixed.
So far each conversation is scored on its own. The training history can be taken into account going forward. What the same person showed in earlier exercises feeds the recommendation for the next exercise. The grade of a single conversation stays unchanged. This part is being introduced now.
It does not change the scoring of a single conversation: weighting, levels, and limits stay as described. Earlier conversations feed the recommendation, not the grade.
Two small rules belong with this: the first sentence comes from the model, not from you, so it is not scored. And passages where speech recognition has obviously misunderstood something are left out.
70%
Goals of this scenario
30%
Competencies
Set for each scenario: the situation, what counts as achieved, and how heavily it weighs.
A fixed set for each training area.
Every individual score uses the same scale
8–10
shown reliably
6–7
visible, not consistent
4–5
partial
0–3
barely shown
The competencies depend on the training area — what matters in a negotiation is different from what matters in a leadership conversation.
04
A level that only says “good” means something different to everyone.
When the behavior is described concretely, everyone shares the same reference point. The room for interpretation gets smaller.
That is why each competency is described across four written levels. We do not score who someone is. We score what was visible in this situation. Two examples from different training areas:
Example 1 · Leadership · “active listening”
8–10 Asks targeted questions and restates what came across.
Picks up earlier points later. The other person feels understood.
6–7 Listens and asks follow-ups, but not throughout.
Some signals from the other person go unaddressed.
4–5 Asks questions but barely hears the answer.
Their own agenda drives the conversation.
0–3 Talks alone for long stretches or interrupts.
No dialogue forms.
Example 2 · Sales · “handling objections”
8–10 Acknowledges the objection and asks what is behind it.
Then addresses that point directly. The conversation continues.
6–7 Responds to the objection, but stays on the surface.
The underlying reason stays unclear.
4–5 Justifies themselves or repeats their own argument.
The objection is still sitting there afterwards.
0–3 Skips the objection or gives in immediately.
The substance is never discussed.
Illustrative only. Each training area has its own competencies and its own levels — and all of them can be adapted to your organization.
Every judgment is tied to a passage from the conversation. That makes it possible to see what it refers to — and to disagree, if needed.
Active listening
6 of 10
Passage from your conversation
“I understand. Let’s still stick with Friday — I need the numbers by then.”
Two sentences earlier your counterpart had said that several things were happening at once. You heard it and acknowledged it, then went straight back to your deadline. A short follow-up question would have been enough to learn what it actually depended on — and whether Friday was realistic.
An example, reconstructed from a leadership scenario. Quotes always come from the person’s own conversation; invented evidence is not allowed for the evaluating model.
05
We made each of them on purpose. None is a side effect of the technology.
Every piece of feedback is tied to a specific moment in your conversation and describes what happened there — not who you are.
Feedback works while attention stays on the task. When it shifts to the person, the effect fades. In a third of the cases studied, feedback even reduced performance.
Kluger & DeNisi, 1996
Four written levels per competency instead of abstract grades. Each one states what was visible.
Scales defined by observable behavior give every rater the same reference point.
Smith & Kendall, 1963
Every scenario has its own goals: a concrete situation in which it is clear what counts as achieved. Each goal is weighted and scored on its own.
Scoring is based on concrete situations from everyday work, not on general traits.
Flanagan, 1954
The evaluating model assigns individual scores, but not an overall grade. The weighting is fixed in the system.
It is always possible to see how a grade was produced — and it is produced the same way across conversations, groups, and time.
Design decision
Four constraints keep a rating from tipping — neither into harshness nor into vagueness.
3 points
The maximum deduction for unfavorable conversation patterns — such as giving in too early, getting personal instead of staying on the issue, or leaving without a firm outcome. A conversation never loses more than that.
3 contributions
That many intelligible contributions are required at minimum. Below that there is no grade at all: a rating on a thin basis is worse than none.
70/30
Scenario goals count for 70 percent, competencies for 30. The model cannot change that — which is why evaluations can be compared with one another.
2 models
One talks with you, the other evaluates. Whoever played the role does not judge the scene.
Feedback that, after one mistake, only shows the failure no longer tells anyone what to do differently next time.
06
The foundation of the loop is the part you fill.
Do you work with your own conversation guide, leadership principles, or competency model? After we agree on it, we place it into the evaluation. Participants then receive feedback in your language and against your standards, rather than generic ones.
What you bring
Your leadership principles
or competency model
Your conversation guide
if you use one
Real situations
from the everyday work of your teams
What we set up with it
Competencies use your terms
Written the way people in your organization talk about conversations.
Each scenario gets its own goals
Drawn from situations that actually occur for you.
Your model becomes the standard
Scoring follows your framework, not someone else’s.
Tone and boundaries fit you
You supply the material and the terms. We do the setup.
If you bring nothing of your own, that is fine too. The scenarios work out of the box without a particular model.
Negotiation training is an exception. Terms such as best alternative, anchoring, and concession are fixed parts of the competencies — without them a negotiation cannot be scored in a meaningful way.
The 70 to 30 weighting cannot be changed. That is intentional: only then do evaluations stay comparable across teams and across months.
07
Four things an evaluation cannot do.
We name them because they matter more for interpretation than any strength we could list.
Not a suitability assessment.
The rating is not a test and is not a basis for personnel decisions. It is a training instrument.
No validated comparison with human raters.
Our scale is built on the principle of behaviorally anchored rating. We have not yet measured how closely it matches the judgment of experienced trainers.
No individual results for managers.
The result belongs to the person who practiced. Reporting upward happens only at group level — including the patterns from the loop.
No result that repeats word for word.
The evaluation runs with very little randomness and in a fixed format. The same conversation scored twice yields the same judgment, but not the same text word for word.
You can test the second point in live use: two experienced people score a sample of transcripts against the same levels, without seeing the system’s scores. The comparison shows where the two judgments diverge. The effort is about a day, and the result is a defensible number instead of an assumption. Talk to us if you want to plan that for your program.
Smith & Kendall (1963)
Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales.
Journal of Applied Psychology, 47(2), 149–155.
Foundational work on behaviorally anchored rating scales. Evidence that they reduce rating bias is positive, but not consistent.
Kluger & DeNisi (1996)
The effects of feedback interventions on performance.
Psychological Bulletin, 119(2), 254–284.
Meta-analysis of 607 effect sizes. Feedback improves performance on average, but reduces it in more than a third of the cases studied.
Flanagan (1954)
The critical incident technique.
Psychological Bulletin, 51(4), 327–358.
A method for deriving rating criteria from concrete professional situations.
Questions about the methodology? We are glad to answer them in detail — including for works councils, data protection, or specialist teams.
Careertrainer.ai[email protected]As of September 2026