Methodology · AI coaching

The Careertrainer Loop

An AI role-play is a conversation. AI coaching is a system that decides which conversation comes next. The Careertrainer Loop is the model behind it: five stations, a foundation, a yield. This page explains how it works, how the evaluation inside it operates, and where the limits are.

YIELDPatterns over timeThe exit leads back to the entranceENTRANCEOccasion

What is coming up?

STATION 1Pass

How does it actually go?

STATION 2Insight

What was it down to?

STATION 3Repetition

What is different this time?

EXITTransfer

What do you take with you?

The inner circle can run as often as neededFOUNDATIONContext

Company context · product or guideline · persona — without these three layers every scenario stays generic.

The loop at a glance. Each station is described on its own below.

01

What separates the loop from a practice catalog

Three properties that turn separate conversations into coaching.

A catalog provides scenarios and leaves the choice to chance or preference. A coach decides what comes next — and explains why.

  1. 1

    Memory across weeks

    Not only within one conversation, but across many — and across different counterparts of the same type. Only then can a pattern be recognized at all.

  2. 2

    A decision instead of a choice

    A catalog asks what you want to practice. The loop says what is due and why — based on what became visible in recent conversations.

  3. 3

    Tied to an appointment

    Not “practice is good”, but: “Thursday proposal meeting, procurement is blocking on price.” That turns an intention into preparation with a deadline.

02

The seven elements

Five stations, a foundation, a yield.

Entrance

Occasion

What is coming up?

A real appointment, in one sentence. That turns a good intention into preparation with a deadline — and a vitamin into a painkiller.

Station 1

Pass

How does it actually go?

The live conversation. Two separate AI systems: one plays, one scores. The character does not know the scoring standard and therefore cannot be gamed. During the conversation only one sentence is on screen: the goal.

Station 2

Insight

What was it down to?

Self-assessment before the result, then the evaluation with quotes from your own transcript — and the gap between the two. Disagreement is intended: an evaluation you cannot argue with is a verdict.

How it works in detail

Station 3

Repetition

What is different this time?

Not a second attempt at the same scenario, but a named variant: tighter, harder, or transferred to another person. The change is fixed in advance.

Exit

Transfer

What do you take with you?

One sentence before the real appointment — and afterwards the question of whether the assumption held. That closes the circle back to the occasion.

Foundation

Context

What makes the conversation yours?

Company context, product or guideline, persona. Without these three layers every scenario stays generic. The part that grows more valuable over time: after months, objections from real conversations sit there.

Yield

Pattern

What shows up over time?

Not a score, but a statement about recurring behavior toward a certain type of counterpart. Aggregated at team level, that becomes the answer to falling back into the old pattern.

What a pattern sounds like

“With quiet market leaders you give in on the third follow-up. With time-pressure types you do not.”

An illustrative wording, not an evaluation of a real person. Aggregated at team level, the same statement might read: “Your team wins against the price-driven buyer and loses against the waiting decision-maker.”

03

Station 2 in detail: the evaluation

Whoever plays the role does not score it.

During the conversation, one model takes the role of your counterpart. Afterwards, a second, independent model reads the transcript and scores it. The model that played the role does not evaluate the conversation itself.

The evaluating model is not allowed to assign an overall grade. It scores each goal and each competency separately. The system then calculates the overall grade, using a weighting that is the same for every conversation.

What it measures against

The model does not invent its own standards. What should be achieved in a scenario, and how that shows up, is set when the scenario is created — by people, before the conversation. After that it stays fixed.

What the evaluating model receives

  • The transcript of the conversation
  • The scenario goals, with their weighting
  • For each goal, the instruction for what counts as achieved
  • The competencies of the training area, with their four levels
  • The starting situation and the description of the character
  • Earlier conversations by the same person, in preparation

What it does not receive

  • The audio recording — only the text is evaluated
  • Conversations by other people
  • Information about role, department, or tenure
  • The authority to assign an overall grade

What changes once there is memory

So far each conversation is scored on its own. With the AI coach, the history is added: the evaluation sees what the same person showed in earlier conversations and can name a pattern instead of only a snapshot. That is the prerequisite for the stations Repetition and Transfer — and the part of the loop that is still being rolled out at the time of this page.

It does not change the scoring of a single conversation: weighting, levels, and limits stay as described. Earlier conversations feed the recommendation, not the grade.

Two small rules belong with this: the first sentence comes from the model, not from you, so it is not scored. And passages where speech recognition has obviously misunderstood something are left out.

What your grade is made of

70%

Goals of this scenario

30%

Competencies

Set for each scenario: the situation, what counts as achieved, and how heavily it weighs.

A fixed set for each training area.

Every individual score uses the same scale

  • 8–10

    shown reliably

  • 6–7

    visible, not consistent

  • 4–5

    partial

  • 0–3

    barely shown

The competencies depend on the training area — what matters in a negotiation is different from what matters in a leadership conversation.

04

Describing behavior instead of giving school grades

A level that only says “good” means something different to everyone.

When the behavior is described concretely, everyone shares the same reference point. The room for interpretation gets smaller.

That is why each competency is described across four written levels. We do not score who someone is. We score what was visible in this situation. Two examples from different training areas:

Example 1 · Leadership · “active listening”

  1. 8–10 Asks targeted questions and restates what came across.

    Picks up earlier points later. The other person feels understood.

  2. 6–7 Listens and asks follow-ups, but not throughout.

    Some signals from the other person go unaddressed.

  3. 4–5 Asks questions but barely hears the answer.

    Their own agenda drives the conversation.

  4. 0–3 Talks alone for long stretches or interrupts.

    No dialogue forms.

Example 2 · Sales · “handling objections”

  1. 8–10 Acknowledges the objection and asks what is behind it.

    Then addresses that point directly. The conversation continues.

  2. 6–7 Responds to the objection, but stays on the surface.

    The underlying reason stays unclear.

  3. 4–5 Justifies themselves or repeats their own argument.

    The objection is still sitting there afterwards.

  4. 0–3 Skips the objection or gives in immediately.

    The substance is never discussed.

Illustrative only. Each training area has its own competencies and its own levels — and all of them can be adapted to your organization.

What a piece of feedback looks like

Every judgment is tied to a passage from the conversation. That makes it possible to see what it refers to — and to disagree, if needed.

Active listening

6 of 10

Passage from your conversation

“I understand. Let’s still stick with Friday — I need the numbers by then.”

Two sentences earlier your counterpart had said that several things were happening at once. You heard it and acknowledged it, then went straight back to your deadline. A short follow-up question would have been enough to learn what it actually depended on — and whether Friday was realistic.

An example, reconstructed from a leadership scenario. Quotes always come from the person’s own conversation; invented evidence is not allowed for the evaluating model.

05

Four decisions and their basis

We made each of them on purpose. None is a side effect of the technology.

Behavior, not the person

Every piece of feedback is tied to a specific moment in your conversation and describes what happened there — not who you are.

Feedback works while attention stays on the task. When it shifts to the person, the effect fades. In a third of the cases studied, feedback even reduced performance.

Kluger & DeNisi, 1996

Described levels

Four written levels per competency instead of abstract grades. Each one states what was visible.

Scales defined by observable behavior give every rater the same reference point.

Smith & Kendall, 1963

Concrete situations

Every scenario has its own goals: a concrete situation in which it is clear what counts as achieved. Each goal is weighted and scored on its own.

Scoring is based on concrete situations from everyday work, not on general traits.

Flanagan, 1954

The system calculates the overall grade

The evaluating model assigns individual scores, but not an overall grade. The weighting is fixed in the system.

It is always possible to see how a grade was produced — and it is produced the same way across conversations, groups, and time.

Design decision

Rules that always apply

Four constraints keep a rating from tipping — neither into harshness nor into vagueness.

  • 3 points

    The maximum deduction for unfavorable conversation patterns — such as giving in too early, getting personal instead of staying on the issue, or leaving without a firm outcome. A conversation never loses more than that.

  • 3 contributions

    That many intelligible contributions are required at minimum. Below that there is no grade at all: a rating on a thin basis is worse than none.

  • 70/30

    Scenario goals count for 70 percent, competencies for 30. The model cannot change that — which is why evaluations can be compared with one another.

  • 2 models

    One talks with you, the other evaluates. Whoever played the role does not judge the scene.

Feedback that, after one mistake, only shows the failure no longer tells anyone what to do differently next time.

06

Adapted to your organization

The foundation of the loop is the part you fill.

Do you work with your own conversation guide, leadership principles, or competency model? After we agree on it, we place it into the evaluation. Participants then receive feedback in your language and against your standards, rather than generic ones.

What you bring

  • Your leadership principles

    or competency model

  • Your conversation guide

    if you use one

  • Real situations

    from the everyday work of your teams

What we set up with it

  • Competencies use your terms

    Written the way people in your organization talk about conversations.

  • Each scenario gets its own goals

    Drawn from situations that actually occur for you.

  • Your model becomes the standard

    Scoring follows your framework, not someone else’s.

  • Tone and boundaries fit you

    You supply the material and the terms. We do the setup.

If you bring nothing of your own, that is fine too. The scenarios work out of the box without a particular model.

Negotiation training is an exception. Terms such as best alternative, anchoring, and concession are fixed parts of the competencies — without them a negotiation cannot be scored in a meaningful way.

The 70 to 30 weighting cannot be changed. That is intentional: only then do evaluations stay comparable across teams and across months.

07

Limits of the system

Four things an evaluation cannot do.

We name them because they matter more for interpretation than any strength we could list.

  • Not a suitability assessment.

    The rating is not a test and is not a basis for personnel decisions. It is a training instrument.

  • No validated comparison with human raters.

    Our scale is built on the principle of behaviorally anchored rating. We have not yet measured how closely it matches the judgment of experienced trainers.

  • No individual results for managers.

    The result belongs to the person who practiced. Reporting upward happens only at group level — including the patterns from the loop.

  • No result that repeats word for word.

    The evaluation runs with very little randomness and in a fixed format. The same conversation scored twice yields the same judgment, but not the same text word for word.

What you can check

You can test the second point in live use: two experienced people score a sample of transcripts against the same levels, without seeing the system’s scores. The comparison shows where the two judgments diverge. The effort is about a day, and the result is a defensible number instead of an assumption. Talk to us if you want to plan that for your program.

Sources

  1. Smith & Kendall (1963)

    Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales.

    Journal of Applied Psychology, 47(2), 149–155.

    Foundational work on behaviorally anchored rating scales. Evidence that they reduce rating bias is positive, but not consistent.

  2. Kluger & DeNisi (1996)

    The effects of feedback interventions on performance.

    Psychological Bulletin, 119(2), 254–284.

    Meta-analysis of 607 effect sizes. Feedback improves performance on average, but reduces it in more than a third of the cases studied.

  3. Flanagan (1954)

    The critical incident technique.

    Psychological Bulletin, 51(4), 327–358.

    A method for deriving rating criteria from concrete professional situations.

Questions about the methodology? We are glad to answer them in detail — including for works councils, data protection, or specialist teams.

Careertrainer.ai[email protected]As of September 2026