Methodology · AI coaching

A role-play on its own is not yet training

Conversation simulations create practice. Training depends on what happens before and after: which situation is practiced, how the conversation is reflected on, what should be trained next, and how that reaches everyday work. The Careertrainer Loop connects these steps into a training process.

PATTERNAcross several conversationsWhat was learned feeds the next occasion.1Occasion

Start from a concrete conversation

2Simulation

Play the conversation in a realistic setting

3Reflection

Bring self-assessment and evaluation together

4Repetition

Practice again under changed conditions

5Transfer

Apply what was learned at work

Simulation, reflection, and repetition can run more than once.FOUNDATIONContext

Company, role, and situation: products, guidelines, typical objections.

The training process at a glance. Each step is explained below.

01

From separate role-plays to a training process

A scenario library mainly answers what someone wants to practice today.

The Careertrainer Loop goes further. Earlier conversations and recognized development points can be used to suggest a fitting next exercise. Separate simulations then become training that builds on itself.

  1. 1

    Training history

    Exercises are not treated as fully isolated. Across several conversations it can become visible where strengths and difficulties repeat, including with similar roles.

  2. 2

    Next training step

    A library leaves the choice to participants. A coach can also take into account which situations have already been trained, where it got difficult, and which exercise fits that.

  3. 3

    A concrete occasion

    The starting point is a real conversation, or one that is typical for the role — for example Thursday’s proposal meeting, where procurement is blocking on price. The exercise then has a link to everyday work.

02

How the training process is built

Five steps, company context as the foundation, and patterns as the result of several runs.

1

Occasion

Start from a concrete conversation

The starting point is a real, upcoming, or role-typical conversation. People then practice situations in which they want to improve how they lead the conversation, rather than a scene detached from work.

2

Simulation

Play the conversation in a realistic setting

The participant holds the conversation with an AI persona. Its behavior is tuned to the scenario, the role, and the situation, and it reacts to how the conversation develops. Simulation and evaluation are separate: the persona holds the conversation, and the evaluation follows afterwards from the full transcript.

3

Reflection

Bring self-assessment and evaluation together

Before the evaluation appears, the participant assesses the conversation. Feedback on conversational behavior follows, with concrete passages from the transcript. Self-assessment and evaluation can then be compared. The evaluation is a reasoned reading against defined criteria, not a final verdict.

How it works in detail

4

Repetition

Practice again under changed conditions

The next exercise does not simply repeat the same scenario. Individual conditions change, for example the other person’s reaction, how sharp an objection is, or the personality type. In the next simulation the customer may push harder on price, or an employee may dodge the feedback more. Earlier exercises can be used to suggest what should be trained next.

5

Transfer

Apply what was learned at work

Training does not end with the simulation. Before a real conversation, people can note what they want to apply. Afterwards they can check whether the approach worked and what that means for further exercises.

Foundation

Context

Use the actual work context

A useful simulation needs more than a generic role description. Company, role, and the concrete situation can be taken into account: products, internal guidelines, typical objections, conversation goals, or relevant counterparts. The exercises then sit closer to the conversations people actually have in the company.

Pattern

Pattern

See development beyond a single conversation

One role-play is a snapshot. Across several exercises, recurring strengths and difficulties can become visible — in certain phases, with objections, or with certain kinds of counterparts. These patterns can be used to choose further training steps more deliberately.

What a pattern sounds like

“With quiet market leaders you give in on the third follow-up. With time-pressure types you do not.”

An illustrative wording, not an evaluation of a real person. Aggregated at team level, the same statement might read: “Your team wins against the price-driven buyer and loses against the waiting decision-maker.”

03

How the evaluation works

Simulation and evaluation are separate.

During the simulation, one model takes the role of the counterpart. Afterwards, a second, independent model reads the transcript and evaluates it. The persona that held the conversation does not evaluate it. The role stays in the conversation, and the evaluation refers to the full transcript.

The evaluating model is not allowed to assign an overall grade. It scores each goal and each competency separately. The system then calculates the overall grade, using a weighting that is the same for every conversation.

What it measures against

The model does not invent its own standards. What should be achieved in a scenario, and how that shows up, is set when the scenario is created — by people, before the conversation. After that it stays fixed.

What the evaluating model receives

  • The transcript of the conversation
  • The scenario goals, with their weighting
  • For each goal, the instruction for what counts as achieved
  • The competencies of the training area, with their four levels
  • The starting situation and the description of the character
  • Earlier conversations by the same person, in preparation

What it does not receive

  • The audio recording — only the text is evaluated
  • Conversations by other people
  • Information about role, department, or tenure
  • The authority to assign an overall grade

How earlier conversations feed further training

So far each conversation is scored on its own. The training history can be taken into account going forward. What the same person showed in earlier exercises feeds the recommendation for the next exercise. The grade of a single conversation stays unchanged. This part is being introduced now.

It does not change the scoring of a single conversation: weighting, levels, and limits stay as described. Earlier conversations feed the recommendation, not the grade.

Two small rules belong with this: the first sentence comes from the model, not from you, so it is not scored. And passages where speech recognition has obviously misunderstood something are left out.

What your grade is made of

70%

Goals of this scenario

30%

Competencies

Set for each scenario: the situation, what counts as achieved, and how heavily it weighs.

A fixed set for each training area.

Every individual score uses the same scale

  • 8–10

    shown reliably

  • 6–7

    visible, not consistent

  • 4–5

    partial

  • 0–3

    barely shown

The competencies depend on the training area — what matters in a negotiation is different from what matters in a leadership conversation.

04

Describing behavior instead of giving school grades

A level that only says “good” means something different to everyone.

When the behavior is described concretely, everyone shares the same reference point. The room for interpretation gets smaller.

That is why each competency is described across four written levels. We do not score who someone is. We score what was visible in this situation. Two examples from different training areas:

Example 1 · Leadership · “active listening”

  1. 8–10 Asks targeted questions and restates what came across.

    Picks up earlier points later. The other person feels understood.

  2. 6–7 Listens and asks follow-ups, but not throughout.

    Some signals from the other person go unaddressed.

  3. 4–5 Asks questions but barely hears the answer.

    Their own agenda drives the conversation.

  4. 0–3 Talks alone for long stretches or interrupts.

    No dialogue forms.

Example 2 · Sales · “handling objections”

  1. 8–10 Acknowledges the objection and asks what is behind it.

    Then addresses that point directly. The conversation continues.

  2. 6–7 Responds to the objection, but stays on the surface.

    The underlying reason stays unclear.

  3. 4–5 Justifies themselves or repeats their own argument.

    The objection is still sitting there afterwards.

  4. 0–3 Skips the objection or gives in immediately.

    The substance is never discussed.

Illustrative only. Each training area has its own competencies and its own levels — and all of them can be adapted to your organization.

What a piece of feedback looks like

Every judgment is tied to a passage from the conversation. That makes it possible to see what it refers to — and to disagree, if needed.

Active listening

6 of 10

Passage from your conversation

“I understand. Let’s still stick with Friday — I need the numbers by then.”

Two sentences earlier your counterpart had said that several things were happening at once. You heard it and acknowledged it, then went straight back to your deadline. A short follow-up question would have been enough to learn what it actually depended on — and whether Friday was realistic.

An example, reconstructed from a leadership scenario. Quotes always come from the person’s own conversation; invented evidence is not allowed for the evaluating model.

05

Four decisions and their basis

We made each of them on purpose. None is a side effect of the technology.

Behavior, not the person

Every piece of feedback is tied to a specific moment in your conversation and describes what happened there — not who you are.

Feedback works while attention stays on the task. When it shifts to the person, the effect fades. In a third of the cases studied, feedback even reduced performance.

Kluger & DeNisi, 1996

Described levels

Four written levels per competency instead of abstract grades. Each one states what was visible.

Scales defined by observable behavior give every rater the same reference point.

Smith & Kendall, 1963

Concrete situations

Every scenario has its own goals: a concrete situation in which it is clear what counts as achieved. Each goal is weighted and scored on its own.

Scoring is based on concrete situations from everyday work, not on general traits.

Flanagan, 1954

The system calculates the overall grade

The evaluating model assigns individual scores, but not an overall grade. The weighting is fixed in the system.

It is always possible to see how a grade was produced — and it is produced the same way across conversations, groups, and time.

Design decision

Rules that always apply

Four constraints keep a rating from tipping — neither into harshness nor into vagueness.

  • 3 points

    The maximum deduction for unfavorable conversation patterns — such as giving in too early, getting personal instead of staying on the issue, or leaving without a firm outcome. A conversation never loses more than that.

  • 3 contributions

    That many intelligible contributions are required at minimum. Below that there is no grade at all: a rating on a thin basis is worse than none.

  • 70/30

    Scenario goals count for 70 percent, competencies for 30. The model cannot change that — which is why evaluations can be compared with one another.

  • 2 models

    One talks with you, the other evaluates. Whoever played the role does not judge the scene.

Feedback that, after one mistake, only shows the failure no longer tells anyone what to do differently next time.

06

Adapted to your organization

The foundation of the loop is the part you fill.

Do you work with your own conversation guide, leadership principles, or competency model? After we agree on it, we place it into the evaluation. Participants then receive feedback in your language and against your standards, rather than generic ones.

What you bring

  • Your leadership principles

    or competency model

  • Your conversation guide

    if you use one

  • Real situations

    from the everyday work of your teams

What we set up with it

  • Competencies use your terms

    Written the way people in your organization talk about conversations.

  • Each scenario gets its own goals

    Drawn from situations that actually occur for you.

  • Your model becomes the standard

    Scoring follows your framework, not someone else’s.

  • Tone and boundaries fit you

    You supply the material and the terms. We do the setup.

If you bring nothing of your own, that is fine too. The scenarios work out of the box without a particular model.

Negotiation training is an exception. Terms such as best alternative, anchoring, and concession are fixed parts of the competencies — without them a negotiation cannot be scored in a meaningful way.

The 70 to 30 weighting cannot be changed. That is intentional: only then do evaluations stay comparable across teams and across months.

07

Limits of the system

Four things an evaluation cannot do.

We name them because they matter more for interpretation than any strength we could list.

  • Not a suitability assessment.

    The rating is not a test and is not a basis for personnel decisions. It is a training instrument.

  • No validated comparison with human raters.

    Our scale is built on the principle of behaviorally anchored rating. We have not yet measured how closely it matches the judgment of experienced trainers.

  • No individual results for managers.

    The result belongs to the person who practiced. Reporting upward happens only at group level — including the patterns from the loop.

  • No result that repeats word for word.

    The evaluation runs with very little randomness and in a fixed format. The same conversation scored twice yields the same judgment, but not the same text word for word.

What you can check

You can test the second point in live use: two experienced people score a sample of transcripts against the same levels, without seeing the system’s scores. The comparison shows where the two judgments diverge. The effort is about a day, and the result is a defensible number instead of an assumption. Talk to us if you want to plan that for your program.

Sources

  1. Smith & Kendall (1963)

    Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales.

    Journal of Applied Psychology, 47(2), 149–155.

    Foundational work on behaviorally anchored rating scales. Evidence that they reduce rating bias is positive, but not consistent.

  2. Kluger & DeNisi (1996)

    The effects of feedback interventions on performance.

    Psychological Bulletin, 119(2), 254–284.

    Meta-analysis of 607 effect sizes. Feedback improves performance on average, but reduces it in more than a third of the cases studied.

  3. Flanagan (1954)

    The critical incident technique.

    Psychological Bulletin, 51(4), 327–358.

    A method for deriving rating criteria from concrete professional situations.

Questions about the methodology? We are glad to answer them in detail — including for works councils, data protection, or specialist teams.

Careertrainer.ai[email protected]As of September 2026