---
title: "Canary AI QA: Test Defect Detection, Not Benchmark Scores"
canonical: https://wavect.io/blog/canary-ai-qa-defect-detection/
language: en
description: "Evaluate Canary alongside Claude Code or Codex with known defects, clean controls and independent checks. Understand what QA-Bench v0 actually measures."
image: "https://wavect.io/img/blog/headers/header_canary-ai-qa-defect-detection.png"
---

[**Back**](/blog/overview/)

[![Kevin Riedl](/img/team/kevin.webp)](/team/kevin-riedl/)

[Kevin Riedl](/team/kevin-riedl/) https://linkedin.com/in/wsdt

4 min read · 8 October 2026 Last reviewed October 8, 2026

[**Next**](/blog/greptile-base-plus-apex-review-budget/)

# Canary AI QA: Test Defect Detection, Not Benchmark Scores

TL;DR

Canary's QA-Bench v0 evaluates generated verification outputs using an LLM judge. Its score is not an executed bug-detection percentage. Assess adoption with known defects, clean controls, actual test execution and independent reproduction of findings.

**Evidence:** Documentation reviewed on 8 October 2026. This is a researched implementation guide. The pilot below is proposed; we have not run these vendor evaluations or measured their performance.

## What does Canary's QA benchmark prove?

Canary's [QA-Bench v0 methodology](https://www.runcanary.ai/blog/qa-bench-v0) evaluates verification outputs across 35 PRs in four repositories. It scores relevance, coverage and coherence using an LLM judge. Its limitations acknowledge that high-level plans and concrete test scripts differ, and that executed pass/fail tests would be a fairer comparison.

An overall score such as 83.1 therefore does not mean “Canary catches 83.1% of bugs.” It also does not establish how a current Claude Code or Codex configuration performs on your application. Use the benchmark to understand the proposed evaluation problem, then measure the outcome you actually need.

## What should an adoption pilot measure instead?

For each known defect, ask whether the generated and executed test fails on the defective version and passes on the repaired version for the right reason. Separate identifying a flow, writing a test, executing it and reproducing a defect. A good plan can be valuable without being an executed regression check.

Canary's [published product setup reference](https://www.runcanary.ai/#install) is the starting point for current integration. Pin the version and supported workflow you actually test; do not infer compatibility from a model name in an older benchmark. Keep the agent writing the feature separate from the acceptance evidence where possible.

## Which defects belong in the test application?

Use a disposable two-tenant application with a known-good reference version. Introduce one defect at a time, keep defect labels hidden from the evaluated agent and include clean control changes. The cases below are a proposed corpus, not findings from a Canary run.

| Defect | Required assertion | Useful negative control |
| --- | --- | --- |
| Missing tenant ownership check | A cannot read B's record | A can still read its own record |
| Duplicate submission | One logical request creates one record | Two different requests create two records |
| Authorization only in UI | Direct backend write is denied | Authorized role can write |
| Failed step reported as success | User sees failure and stored state agrees | Happy path persists correct state |
| Destructive scope too broad | Delete affects only selected records | Unselected records remain readable |
| Expired session | No privileged write after expiry | Valid session remains functional |

For an intentional security defect, expose only local or isolated synthetic data. Do not introduce it into a shared production environment to obtain a realistic result.

## How do you score the results without fooling yourself?

1. Run the same starting application, PR context and time budget for each evaluated setup.
2. Save plans, generated tests, executed commands, logs and destination state.
3. Reproduce every claimed defect independently. Label infrastructure failures separately.
4. Run each candidate test against both defective and repaired versions.
5. Count confirmed detections, missed seeded defects, false positives on clean controls and reviewer minutes.

Report detection per defect class and sample size. A test that fails because the app never started is not a successful detection. A test that also fails after repair does not establish a useful regression assertion. Keep repeat runs so the result shows variability rather than one lucky attempt.

## Can Canary replace the application's test suite?

Do not make that decision from a vendor score. Preserve deterministic checks for business invariants and critical integrations. Generated exploratory tests can expand coverage or discover cases your suite lacks. Promote a useful generated test into the maintained suite only after reviewing its assertions, fixture ownership and stability.

Keep an explicit gate for permission leaks and destructive errors. A high average score should not compensate for missing a critical tenant-isolation defect. Define the stop condition before you look at results.

## How should you budget AI QA?

Measure cost per accepted verification, including execution infrastructure, repeated runs and human triage. Count duplicate reports as one defect. Track how many tests remain useful after repair and after the next product change. A low generation price can become expensive when every PR requires minutes of false-positive investigation.

## When should you add Canary to a coding-agent workflow?

When an isolated pilot shows useful additional detections or maintained regression tests at acceptable triage cost. If the value is primarily flow planning, adopt it for that job and label it accordingly. For a production decision, bring your defect corpus and acceptance rules to [scope an independent QA evaluation](/contact/).

[Download the proposed pilot protocol (JSON). It contains acceptance cases and empty result fields, not measured vendor results.](/downloads/canary-ai-qa-defect-detection-pilot.json)

## Related implementation guidance

[Greptile Base vs Plus vs Apex: A PR Review Budget](/blog/greptile-base-plus-apex-review-budget/). [Arga Labs vs Archal: Stateful Agent Integration Tests](/blog/arga-vs-archal-agent-integration-testing/).

## Sources checked

- [Canary: QA-Bench v0](https://www.runcanary.ai/blog/qa-bench-v0)
- [Canary: CLI](https://www.runcanary.ai/#install)

**Independence and trademarks:** Wavect publishes this page and is itself a provider, so we have a commercial interest in it. We are not affiliated with, endorsed by or partnered with the other companies named here, and all third-party company names, brands and trademarks are the property of their respective owners. Statements about other providers are taken from publicly available sources, primarily their own published pages, as of the review date shown on this page, and may have changed since. Please verify them directly before you decide. This page was written to the best of our knowledge and with the intent to remain objective. If you believe anything here is inaccurate or unfair, write to us and we will correct it: [office@wavect.io](mailto:office@wavect.io)

QA and production readiness

## Continue through this cluster

Testing, audits, maintenance and hardening practices for reliable production software.

[Start with the cornerstone**QA for AI-Generated Code**](/blog/qa-for-ai-generated-code/)

- [Greptile Base vs Plus vs Apex: A PR Review Budget](/blog/greptile-base-plus-apex-review-budget/)
- [Arga Labs vs Archal: Stateful Agent Integration Tests](/blog/arga-vs-archal-agent-integration-testing/)
- [Cua for Desktop QA: A Browser-to-Native Test Protocol](/blog/cua-desktop-qa-browser-native-workflow/)
- [Browser Use vs Playwright: Verify Authenticated Actions After Timeouts](/blog/browser-use-vs-playwright-authenticated-workflow/)
- [ChatGPT Dots + GitHub: From Bug Report to Reviewed PR](/blog/chatgpt-dots-github-bug-triage/)

[**Back**](/blog/overview/)

[![Kevin Riedl](/img/team/kevin.webp)](/team/kevin-riedl/)

[Kevin Riedl](/team/kevin-riedl/) https://linkedin.com/in/wsdt

4 min read · 8 October 2026 Last reviewed October 8, 2026

[**Next**](/blog/greptile-base-plus-apex-review-budget/)

## Structured Data

```json
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@id": "https://wavect.io/#organization",
      "@type": [
        "Organization",
        "ProfessionalService",
        "LocalBusiness"
      ],
      "employee": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "founder": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "legalRepresentative": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "name": "Wavect GmbH",
      "subjectOf": {
        "@id": "https://wavect.io/verified-claims.json#dataset",
        "@type": "Dataset",
        "creator": {
          "@id": "https://wavect.io/#organization",
          "@type": [
            "Organization",
            "ProfessionalService",
            "LocalBusiness"
          ]
        },
        "description": "A machine-readable registry of quantitative and qualitative claims published by Wavect, with review dates, localized page appearances and public third-party citations where available.",
        "inLanguage": "en",
        "isAccessibleForFree": true,
        "license": "https://creativecommons.org/licenses/by/4.0/",
        "name": "Wavect verified publication claims",
        "url": "https://wavect.io/verified-claims.json"
      },
      "url": "https://wavect.io/"
    },
    {
      "@id": "https://wavect.io/team/kevin-riedl/#person",
      "@type": "Person",
      "jobTitle": "Managing Director",
      "name": "Kevin Riedl",
      "sameAs": [
        "https://www.wikidata.org/wiki/Q139796365",
        "https://www.linkedin.com/in/wsdt",
        "https://github.com/wsdt"
      ],
      "url": "https://wavect.io/team/kevin-riedl/",
      "worksFor": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      }
    },
    {
      "@id": "https://wavect.io/team/christof-jori/#person",
      "@type": "Person",
      "jobTitle": "Managing Director",
      "name": "Christof Jori",
      "sameAs": [
        "https://www.wikidata.org/wiki/Q139796367",
        "https://www.linkedin.com/in/jocr77/",
        "https://github.com/jo-chris"
      ],
      "url": "https://wavect.io/team/christof-jori/",
      "worksFor": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      }
    },
    {
      "@id": "https://wavect.io/#website",
      "@type": "WebSite",
      "inLanguage": [
        "en",
        "de",
        "es",
        "zh"
      ],
      "name": "Wavect",
      "potentialAction": {
        "@type": "SearchAction",
        "query-input": "required name=search_term_string",
        "target": {
          "@type": "EntryPoint",
          "urlTemplate": "https://wavect.io/search/?q={search_term_string}"
        }
      },
      "publisher": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      },
      "url": "https://wavect.io/"
    },
    {
      "@id": "https://wavect.io/blog/canary-ai-qa-defect-detection/#webpage",
      "@type": "WebPage",
      "dateModified": "2026-10-08",
      "inLanguage": "en",
      "isPartOf": {
        "@id": "https://wavect.io/#website",
        "@type": "WebSite"
      },
      "lastReviewed": "2026-10-08",
      "url": "https://wavect.io/blog/canary-ai-qa-defect-detection/"
    }
  ]
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  "abstract": "Canary's QA-Bench v0 evaluates generated verification outputs using an LLM judge. Its score is not an executed bug-detection percentage. Assess adoption with known defects, clean controls, actual test execution and independent reproduction of findings.",
  "articleBody": " Blog overview/Delivery and QA/QA and production readiness Canary AI QA: Test Defect Detection, Not Benchmark Scores TL;DR Canary's QA-Bench v0 evaluates generated verification outputs using an LLM judge. Its score is not an executed bug-detection percentage. Assess adoption with known defects, clean controls, actual test execution and independent reproduction of findings. Evidence: Documentation reviewed on 8 October 2026. This is a researched implementation guide. The pilot below is proposed; we have not run these vendor evaluations or measured their performance. What does Canary's QA benchmark prove? Canary's QA-Bench v0 methodology evaluates verification outputs across 35 PRs in four repositories. It scores relevance, coverage and coherence using an LLM judge. Its limitations acknowledge that high-level plans and concrete test scripts differ, and that executed pass/fail tests would be a fairer comparison. An overall score such as 83.1 therefore does not mean “Canary catches 83.1% of bugs.” It also does not establish how a current Claude Code or Codex configuration performs on your application. Use the benchmark to understand the proposed evaluation problem, then measure the outcome you actually need. What should an adoption pilot measure instead? For each known defect, ask whether the generated and executed test fails on the defective version and passes on the repaired version for the right reason. Separate identifying a flow, writing a test, executing it and reproducing a defect. A good plan can be valuable without being an executed regression check. Canary's published product setup reference is the starting point for current integration. Pin the version and supported workflow you actually test; do not infer compatibility from a model name in an older benchmark. Keep the agent writing the feature separate from the acceptance evidence where possible. Which defects belong in the test application? Use a disposable two-tenant application with a known-good reference version. Introduce one defect at a time, keep defect labels hidden from the evaluated agent and include clean control changes. The cases below are a proposed corpus, not findings from a Canary run. DefectRequired assertionUseful negative control Missing tenant ownership checkA cannot read B's recordA can still read its own recordDuplicate submissionOne logical request creates one recordTwo different requests create two recordsAuthorization only in UIDirect backend write is deniedAuthorized role can writeFailed step reported as successUser sees failure and stored state agreesHappy path persists correct stateDestructive scope too broadDelete affects only selected recordsUnselected records remain readableExpired sessionNo privileged write after expiryValid session remains functional For an intentional security defect, expose only local or isolated synthetic data. Do not introduce it into a shared production environment to obtain a realistic result. How do you score the results without fooling yourself? Run the same starting application, PR context and time budget for each evaluated setup.Save plans, generated tests, executed commands, logs and destination state.Reproduce every claimed defect independently. Label infrastructure failures separately.Run each candidate test against both defective and repaired versions.Count confirmed detections, missed seeded defects, false positives on clean controls and reviewer minutes. Report detection per defect class and sample size. A test that fails because the app never started is not a successful detection. A test that also fails after repair does not establish a useful regression assertion. Keep repeat runs so the result shows variability rather than one lucky attempt. Can Canary replace the application's test suite? Do not make that decision from a vendor score. Preserve deterministic checks for business invariants and critical integrations. Generated exploratory tests can expand coverage or discover cases your suite lacks. Promote a useful generated test into the maintained suite only after reviewing its assertions, fixture ownership and stability. Keep an explicit gate for permission leaks and destructive errors. A high average score should not compensate for missing a critical tenant-isolation defect. Define the stop condition before you look at results. How should you budget AI QA? Measure cost per accepted verification, including execution infrastructure, repeated runs and human triage. Count duplicate reports as one defect. Track how many tests remain useful after repair and after the next product change. A low generation price can become expensive when every PR requires minutes of false-positive investigation. When should you add Canary to a coding-agent workflow? When an isolated pilot shows useful additional detections or maintained regression tests at acceptable triage cost. If the value is primarily flow planning, adopt it for that job and label it accordingly. For a production decision, bring",
  "articleSection": "Engineering",
  "author": {
    "@id": "https://wavect.io/team/kevin-riedl/#person",
    "@type": "Person",
    "name": "Kevin Riedl",
    "sameAs": [
      "https://www.wikidata.org/wiki/Q139796365",
      "https://www.linkedin.com/in/wsdt",
      "https://github.com/wsdt"
    ],
    "url": "https://wavect.io/team/kevin-riedl/"
  },
  "citation": [
    {
      "@type": "WebPage",
      "name": "QA-Bench v0 methodology",
      "url": "https://www.runcanary.ai/blog/qa-bench-v0"
    },
    {
      "@type": "WebPage",
      "name": "published product setup reference",
      "url": "https://www.runcanary.ai/#install"
    }
  ],
  "dateModified": "2026-10-08",
  "datePublished": "2026-10-08",
  "description": "Canary's QA-Bench v0 evaluates generated verification outputs using an LLM judge. Its score is not an executed bug-detection percentage. Assess adoption with known defects, clean controls, actual test execution and independent reproduction of findings.",
  "headline": "Canary AI QA: Test Defect Detection, Not Benchmark Scores",
  "image": "https://wavect.io/img/blog/headers/header_canary-ai-qa-defect-detection.svg",
  "inLanguage": "en",
  "keywords": "Engineering, AI agents",
  "mainEntityOfPage": {
    "@id": "https://wavect.io/blog/canary-ai-qa-defect-detection/",
    "@type": "WebPage"
  },
  "publisher": {
    "@id": "https://wavect.io/#organization",
    "@type": [
      "Organization",
      "ProfessionalService",
      "LocalBusiness"
    ]
  },
  "url": "https://wavect.io/blog/canary-ai-qa-defect-detection/",
  "wordCount": 1130
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "BreadcrumbList",
  "itemListElement": [
    {
      "@type": "ListItem",
      "item": "https://wavect.io/",
      "name": "Home",
      "position": 1
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/overview/",
      "name": "Blog overview",
      "position": 2
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/topics/delivery-qa/",
      "name": "Delivery and QA",
      "position": 3
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/clusters/qa-production/",
      "name": "QA and production readiness",
      "position": 4
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/canary-ai-qa-defect-detection/",
      "name": "Canary AI QA: Test Defect Detection, Not Benchmark Scores",
      "position": 5
    }
  ]
}
```
