Quality Assurance (QA)
QA is how you measure and improve agent performance: design scorecards, evaluate calls against them, let agents see and dispute their scores, and keep evaluators consistent with calibration. This chapter covers the full workflow.
In the app: all under the Quality menu — Scorecards, Evaluations, My Evaluations, and Calibration.
1. Scorecards — the evaluation template
A Scorecard is a reusable template of things you grade on a call. It contains sections and weighted criteria.
Scorecards with their criteria count and active/draft status. Build them with weighted criteria, auto-fail, and deterministic auto-QA.
Each criterion has:
- A question (e.g. "Did the agent verify the caller's identity?").
- An answer type — yes/no, 1–5 scale, or N/A.
- A weight — how much it counts toward the total.
- An optional auto-fail flag — failing this one criterion fails the whole evaluation (e.g. a compliance breach), regardless of other scores.
- Optional auto-rules — deterministic checks that pre-fill the answer from call data (see §3).
To create one: Quality → Scorecards → New Scorecard:
The builder: name the scorecard, then add weighted criteria — each with a question, answer type (yes-no / scale / N/A), a weight, and an optional auto-fail flag.
| Field | What it does |
|---|---|
| Name | The scorecard's name (e.g. "Sales QA v1"). |
| Criteria | Each question you grade on. Add as many as you need. |
| Weight | How much each criterion counts toward the total score. |
| Auto-fail | Mark a criterion so that failing it fails the whole evaluation. |
Add your criteria, set answer types and weights, then Save.
Scorecard versioning — ✅ Available
Scorecards are versioned. When you change a published scorecard, past evaluations keep the version they were scored against, so historical results stay accurate and comparable. Edits create a new version rather than rewriting history.
2. Evaluations — grading a call
An Evaluation applies a scorecard to a specific call.
To evaluate a call: open the call in Call Review → Evaluate → pick a scorecard → Create Evaluation (or Quality → Evaluations → open a row). The evaluation opens as a panel of the scorecard's criteria — each with its section, question, and weight, and a Yes / No answer. Answer each while listening, then Submit. Orbit computes the total score (%) and triggers any auto-fail.
An evaluation moves through states: draft → submitted → acknowledged / disputed.
You currently choose calls to evaluate manually — find them in Call Records / Call Review (filter by agent, queue, sentiment, or flags) and click Evaluate. Automated sampling/quota assignment (e.g. "N random evals per agent per week" as a work queue) and QA trend reporting (per-agent score over time, team average, weakest criteria) are planned enhancements — for now, track trends via the Dashboard's sentiment/Top-Users widgets and exported Call Records.
Auto-QA pre-fill — ✅ Available
To save time, Orbit can pre-fill criteria deterministically (no AI guesswork) from call data: transcript keyword matches, sentiment from the call's score, hold count and talk time from the call journey, and greeting/compliance patterns. You review and adjust the pre-filled answers before submitting.
AI Judge — LLM-scored evaluations — ✅ Available
Beyond deterministic rules, the AI Judge reads the call's transcript and answers each criterion for you, with a short reason. Open an evaluation and press AI Judge: every question is pre-filled Yes/No and tagged auto-filled, with the model's rationale shown beneath it (e.g. "AI: No greeting or introduction by the agent.").
Each criterion shows its section, question, weight, and an AI: … rationale from the AI Judge. You review, override any answer, and Submit — or press Re-AI Judge to run it again.
The point is assist, not replace: the AI Judge gives you a fast first pass and its reasoning, but a human always reviews and submits — you can flip any answer before it counts. Use it to triage a large backlog, then spot-check.
A quick evaluation cycle with AI Judge:
- Open the call in Call Review → Evaluate → pick a scorecard → Create Evaluation.
- Press AI Judge — every criterion is pre-filled with a Yes/No and a reason.
- Review each answer against what you hear; override any you disagree with. (Press Re-AI Judge if you changed the scorecard or want a fresh pass.)
- Submit. Orbit computes the score and applies any auto-fail.
- The agent then acknowledges or disputes it (§3).
3. The agent's view — My Evaluations, acknowledge & dispute
Agents see evaluations of their own calls under My Evaluations.
My Evaluations lists each of the agent's scored calls with its scorecard, score, and status — open one to read the criteria and acknowledge or dispute it.
Dispute workflow — ✅ Available
For each evaluation an agent can:
- Acknowledge — confirm they've read and accept it, or
- Dispute — formally contest a score they disagree with, with a reason.
A disputed evaluation is flagged for the supervisor to review and resolve, creating a documented back-and-forth instead of an unappealable score.
Resolving a dispute (supervisor): open the disputed evaluation and record an outcome — Upheld (original score stands), Overturned (the agent's objection is accepted), or Adjusted (you change specific answers/score). The resolution and its reason are stored as an append-only history on the evaluation, so there's a clear audit trail of who changed what and why.
4. Calibration — keeping evaluators consistent — ✅ Available
Calibration checks that your evaluators grade the same call the same way. In a calibration session, multiple supervisors independently score one shared call; Orbit then compares their scores to surface disagreement.
Use it to:
- spot evaluators who are systematically too harsh or too lenient,
- align the team on how each criterion should be interpreted,
- defend the fairness of your QA program.
To run one: Quality → Calibration → New session → choose the call and the evaluators → each scores independently → review the agreement report.
The agreement report shows you where evaluators diverge so you can fix it:
| The report shows | What it tells you |
|---|---|
| Score spread | How far apart the evaluators' totals are on the same call. |
| Outliers | Which evaluator is systematically harsher or more lenient than the group. |
| Per-question divergence | Which specific criteria people interpret differently — exactly what to clarify in your scorecard or training. |
5. A typical QA cycle
- Design a scorecard (§1).
- Evaluate a sample of each agent's calls (§2), letting auto-QA pre-fill what it can.
- Agents acknowledge or dispute their results (§3).
- Periodically calibrate evaluators on a shared call (§4).
- Coach using the scores, sentiment, and call recordings.