Vela support agent, July 2026 run

A benchmark that shows its working.

We test Vela against the same model without our retrieval layer. Each row shows what counts as a pass and how many questions were in the set.

Vela Baseline Scale 0 to 100%. Higher is better.
  • Grounded answers

    Answers that cite an approved help article and need no edits from a reviewer.

    Test set: 1,200 past support questions, anonymized

    94 %

    +23 pts over baseline

  • Policy conflicts caught

    Questions where two help articles disagree, flagged before an answer goes out.

    Test set: 240 seeded conflict cases

    89 %

    +26 pts over baseline

  • Out-of-scope questions handed off

    Billing disputes, legal requests, and account recovery sent to a person.

    Test set: 180 out-of-scope questions

    97 %

    +15 pts over baseline

Method: fixed prompts, frozen help-center snapshot from 1 July 2026, and blind review by three support leads. Disagreements went to a fourth reviewer.