Skip to main content

Quality Assurance (QA)

SupervisorCustomer AdminAgent

QA is how you measure and improve agent performance: design scorecards, evaluate calls against them, let agents see and dispute their scores, and keep evaluators consistent with calibration. This chapter covers the full workflow.

In the app: all under the Quality menu — Scorecards, Evaluations, My Evaluations, and Calibration.


1. Scorecards — the evaluation template

A Scorecard is a reusable template of things you grade on a call. It contains sections and weighted criteria.

The Scorecards screen listing evaluation templates Scorecards with their criteria count and active/draft status. Build them with weighted criteria, auto-fail, and deterministic auto-QA.

Each criterion has:

  • A question (e.g. "Did the agent verify the caller's identity?").
  • An answer typeyes/no, 1–5 scale, or N/A.
  • A weight — how much it counts toward the total.
  • An optional auto-fail flag — failing this one criterion fails the whole evaluation (e.g. a compliance breach), regardless of other scores.
  • Optional auto-rules — deterministic checks that pre-fill the answer from call data (see §3).

To create one: Quality → ScorecardsNew Scorecard:

The Scorecard builder The builder: name the scorecard, then add weighted criteria — each with a question, answer type (yes-no / scale / N/A), a weight, and an optional auto-fail flag.

FieldWhat it does
NameThe scorecard's name (e.g. "Sales QA v1").
CriteriaEach question you grade on. Add as many as you need.
WeightHow much each criterion counts toward the total score.
Auto-failMark a criterion so that failing it fails the whole evaluation.

Add your criteria, set answer types and weights, then Save.

Scorecard versioning — ✅ Available

Scorecards are versioned. When you change a published scorecard, past evaluations keep the version they were scored against, so historical results stay accurate and comparable. Edits create a new version rather than rewriting history.


2. Evaluations — grading a call

An Evaluation applies a scorecard to a specific call.

To evaluate a call: open the call in Call ReviewEvaluate → pick a scorecard → Create Evaluation (or Quality → Evaluations → open a row). The evaluation opens as a panel of the scorecard's criteria — each with its section, question, and weight, and a Yes / No answer. Answer each while listening, then Submit. Orbit computes the total score (%) and triggers any auto-fail.

An evaluation moves through states: draft → submitted → acknowledged / disputed.

Picking calls to evaluate (today vs. planned)

You currently choose calls to evaluate manually — find them in Call Records / Call Review (filter by agent, queue, sentiment, or flags) and click Evaluate. Automated sampling/quota assignment (e.g. "N random evals per agent per week" as a work queue) and QA trend reporting (per-agent score over time, team average, weakest criteria) are planned enhancements — for now, track trends via the Dashboard's sentiment/Top-Users widgets and exported Call Records.

Auto-QA pre-fill — ✅ Available

To save time, Orbit can pre-fill criteria deterministically (no AI guesswork) from call data: transcript keyword matches, sentiment from the call's score, hold count and talk time from the call journey, and greeting/compliance patterns. You review and adjust the pre-filled answers before submitting.

AI Judge — LLM-scored evaluations — ✅ Available

Beyond deterministic rules, the AI Judge reads the call's transcript and answers each criterion for you, with a short reason. Open an evaluation and press AI Judge: every question is pre-filled Yes/No and tagged auto-filled, with the model's rationale shown beneath it (e.g. "AI: No greeting or introduction by the agent.").

An AI-judged evaluation: each criterion pre-filled Yes/No with the AI's reasoning, and Re-AI Judge / Submit Each criterion shows its section, question, weight, and an AI: … rationale from the AI Judge. You review, override any answer, and Submit — or press Re-AI Judge to run it again.

The point is assist, not replace: the AI Judge gives you a fast first pass and its reasoning, but a human always reviews and submits — you can flip any answer before it counts. Use it to triage a large backlog, then spot-check.

A quick evaluation cycle with AI Judge:

  1. Open the call in Call ReviewEvaluate → pick a scorecard → Create Evaluation.
  2. Press AI Judge — every criterion is pre-filled with a Yes/No and a reason.
  3. Review each answer against what you hear; override any you disagree with. (Press Re-AI Judge if you changed the scorecard or want a fresh pass.)
  4. Submit. Orbit computes the score and applies any auto-fail.
  5. The agent then acknowledges or disputes it (§3).

3. The agent's view — My Evaluations, acknowledge & dispute

Agents see evaluations of their own calls under My Evaluations.

The My Evaluations screen where an agent reviews their scored calls My Evaluations lists each of the agent's scored calls with its scorecard, score, and status — open one to read the criteria and acknowledge or dispute it.

Dispute workflow — ✅ Available

For each evaluation an agent can:

  • Acknowledge — confirm they've read and accept it, or
  • Dispute — formally contest a score they disagree with, with a reason.

A disputed evaluation is flagged for the supervisor to review and resolve, creating a documented back-and-forth instead of an unappealable score.

Resolving a dispute (supervisor): open the disputed evaluation and record an outcome — Upheld (original score stands), Overturned (the agent's objection is accepted), or Adjusted (you change specific answers/score). The resolution and its reason are stored as an append-only history on the evaluation, so there's a clear audit trail of who changed what and why.


4. Calibration — keeping evaluators consistent — ✅ Available

Calibration checks that your evaluators grade the same call the same way. In a calibration session, multiple supervisors independently score one shared call; Orbit then compares their scores to surface disagreement.

Use it to:

  • spot evaluators who are systematically too harsh or too lenient,
  • align the team on how each criterion should be interpreted,
  • defend the fairness of your QA program.

To run one: Quality → Calibration → New session → choose the call and the evaluators → each scores independently → review the agreement report.

The agreement report shows you where evaluators diverge so you can fix it:

The report showsWhat it tells you
Score spreadHow far apart the evaluators' totals are on the same call.
OutliersWhich evaluator is systematically harsher or more lenient than the group.
Per-question divergenceWhich specific criteria people interpret differently — exactly what to clarify in your scorecard or training.

5. A typical QA cycle

  1. Design a scorecard (§1).
  2. Evaluate a sample of each agent's calls (§2), letting auto-QA pre-fill what it can.
  3. Agents acknowledge or dispute their results (§3).
  4. Periodically calibrate evaluators on a shared call (§4).
  5. Coach using the scores, sentiment, and call recordings.