---
title: "Arga Labs vs Archal: Stateful Agent Integration Tests"
canonical: https://wavect.io/blog/arga-vs-archal-agent-integration-testing/
language: en
description: "Compare Arga and Archal for repeatable integration tests. Use a GitHub-to-Slack protocol covering duplicate events, interrupted runs and forbidden writes."
image: "https://wavect.io/img/blog/headers/header_arga-vs-archal-agent-integration-testing.png"
---

[**Back**](/blog/overview/)

[![Kevin Riedl](/img/team/kevin.webp)](/team/kevin-riedl/)

[Kevin Riedl](/team/kevin-riedl/) https://linkedin.com/in/wsdt

4 min read · 8 October 2026 Last reviewed October 8, 2026

[**Next**](/blog/agent-harness-engineering/)

# Arga Labs vs Archal: Stateful Agent Integration Tests

TL;DR

Choose Arga or Archal by the exact service operations, starting state and CI lifecycle your workflow needs. Stateful twins expose cross-call effects that simple mocks can miss, but passing a simulation does not verify the real provider. Keep a separate contract smoke test.

**Evidence:** Documentation reviewed on 8 October 2026. This is a researched implementation guide. The pilot below is proposed; we have not run these vendor evaluations or measured their performance.

## When do stateful API environments beat simple mocks?

When correctness depends on what earlier actions changed. A mock returning “success” for posting a Slack message does not prove that a retry created only one message, that an unauthorized channel stayed unchanged, or that a later read saw the correct state. A stateful environment lets you inspect those consequences.

The [YC profile for Arga Labs](https://www.ycombinator.com/companies/arga-labs) identifies the company at `argalabs.com`. The launch describes service replicas and earlier application-staging behavior. Use current documentation to scope a purchase; do not assume a historical launch describes the current deployment contract.

## What is the current Arga versus Archal boundary?

Arga's [current test-workflow documentation](https://docs.argalabs.com/features/validate-modes) separates Twin Runs from browser Test Runs and PR Test Runs. PR Test Runs use the configured application URL; they do not deploy the changed app into a per-PR Arga environment. Your CI must still supply the correct application build and reachable URL.

Archal's [sandbox starting-state documentation](https://docs.archal.ai/sandboxes/starting-state) describes versioned samples and explicit JSON, with guarded SQL for Supabase. State loaded during creation establishes the reset baseline. This differs from older material describing natural-language scenario generation. Compare current state contracts, not product slogans.

| Decision | Arga documentation | Archal documentation |
| --- | --- | --- |
| Test your application's UI | Browser tests against a reachable URL | Provide your own application and test runner |
| Reproduce starting state | Saved scenarios and twin state | Versioned samples or validated explicit state |
| Validate the changed app | Supply the intended deployment URL | Run your app or agent against returned service endpoints |

## What should a GitHub-to-Slack test assert?

Proposed task: when an authorized issue is labeled “ready,” post one summary in the allowed Slack channel and save a durable delivery reference. Prepare the same synthetic repository, issue, channel and permission intent for both products. Confirm the required API operations are actually supported before provisioning. These are test requirements, not claims that both vendors implement every fault.

| Case | Injection | Required final state |
| --- | --- | --- |
| Happy path | One allowed label event | One message and one delivery reference |
| Duplicate event | Deliver the same event twice | Still one logical delivery |
| Interrupted run | Stop after send, before local acknowledgement | Reconcile existing message before retry |
| Denied channel | Remove posting permission | No message; explicit failure |
| Unexpected response | Add a controlled malformed response in the harness | No false success; inspectable error |
| Reset | Restore the fixture and rerun | Same initial records and equivalent outcome |

Check destination records independently of the agent's final answer. “Done” is not a state assertion. Keep an application database fixture too: resetting a vendor twin does not automatically reset your own delivery ledger.

## How should CI isolate credentials and failures?

Use returned sandbox endpoints and scoped credentials, with an allowlist that prevents real-provider fallback. Archal's [current lifecycle and authentication reference](https://docs.archal.ai/llms.txt) separates workspace control credentials from environment credentials. A failed environment startup should fail the job, rather than redirect test traffic to production.

Save a before-state snapshot, request trace, after-state snapshot and operation ledger. In cleanup, tear down the session even when assertions fail. Reset between independent cases, and serialize tests that intentionally share state. Redact credentials and personal data before storing artifacts.

## What do these environments not prove?

Check Arga's [per-twin support and limitations](https://docs.argalabs.com/concepts/twin-reference) for each required endpoint. Simulations can differ from providers in rate limits, OAuth, timing, webhook delivery and undocumented behavior. Maintain a narrow real-provider test account for contract checks, while using twins for repeatable fault and state tests. Passing one should not waive the other.

Archal's [usage documentation](https://docs.archal.ai/pricing-usage) currently bills ready environment duration at USD 0.10 per environment-minute, prorated by seconds. Two ready environments for ten minutes model USD 2, before any other tool costs. This is arithmetic, not a measured run. Include setup, teardown, idle time and the current availability of paid continuation in your adoption discussion. Paid continuation is currently not self-serve in the checked documentation; confirm access before planning recurring CI usage.

## Which should you choose?

Choose the product that supports your exact workflow and can reproduce failures with an affordable, reliable CI lifecycle. Arga is worth evaluating when reusable browser flows and service twins fit together; Archal is worth evaluating when explicit state and provider-shaped API environments fit your existing runner. Bring the operation list and failure protocol to [plan an integration-testing pilot](/contact/).

[Download the proposed pilot protocol (JSON). It contains acceptance cases and empty result fields, not measured vendor results.](/downloads/arga-vs-archal-agent-integration-testing-pilot.json)

## Related implementation guidance

[AI Agent Harness, Explained: The Reliability Layer Around an LLM](/blog/agent-harness-engineering/). [Canary AI QA: Test Defect Detection, Not Benchmark Scores](/blog/canary-ai-qa-defect-detection/).

## Sources checked

- [YC: Arga Labs](https://www.ycombinator.com/companies/arga-labs)
- [Arga: Validate Modes](https://docs.argalabs.com/features/validate-modes)
- [Archal: Starting State](https://docs.archal.ai/sandboxes/starting-state)
- [Archal: llms.txt](https://docs.archal.ai/llms.txt)
- [Arga: Twin Reference](https://docs.argalabs.com/concepts/twin-reference)
- [Archal: Pricing & Usage](https://docs.archal.ai/pricing-usage)

**Independence and trademarks:** Wavect publishes this page and is itself a provider, so we have a commercial interest in it. We are not affiliated with, endorsed by or partnered with the other companies named here, and all third-party company names, brands and trademarks are the property of their respective owners. Statements about other providers are taken from publicly available sources, primarily their own published pages, as of the review date shown on this page, and may have changed since. Please verify them directly before you decide. This page was written to the best of our knowledge and with the intent to remain objective. If you believe anything here is inaccurate or unfair, write to us and we will correct it: [office@wavect.io](mailto:office@wavect.io)

QA and production readiness

## Continue through this cluster

Testing, audits, maintenance and hardening practices for reliable production software.

[Start with the cornerstone**QA for AI-Generated Code**](/blog/qa-for-ai-generated-code/)

- [Greptile Base vs Plus vs Apex: A PR Review Budget](/blog/greptile-base-plus-apex-review-budget/)
- [Canary AI QA: Test Defect Detection, Not Benchmark Scores](/blog/canary-ai-qa-defect-detection/)
- [Cua for Desktop QA: A Browser-to-Native Test Protocol](/blog/cua-desktop-qa-browser-native-workflow/)
- [Browser Use vs Playwright: Verify Authenticated Actions After Timeouts](/blog/browser-use-vs-playwright-authenticated-workflow/)
- [ChatGPT Dots + GitHub: From Bug Report to Reviewed PR](/blog/chatgpt-dots-github-bug-triage/)

[**Back**](/blog/overview/)

[![Kevin Riedl](/img/team/kevin.webp)](/team/kevin-riedl/)

[Kevin Riedl](/team/kevin-riedl/) https://linkedin.com/in/wsdt

4 min read · 8 October 2026 Last reviewed October 8, 2026

[**Next**](/blog/agent-harness-engineering/)

## Structured Data

```json
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@id": "https://wavect.io/#organization",
      "@type": [
        "Organization",
        "ProfessionalService",
        "LocalBusiness"
      ],
      "employee": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "founder": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "legalRepresentative": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "name": "Wavect GmbH",
      "subjectOf": {
        "@id": "https://wavect.io/verified-claims.json#dataset",
        "@type": "Dataset",
        "creator": {
          "@id": "https://wavect.io/#organization",
          "@type": [
            "Organization",
            "ProfessionalService",
            "LocalBusiness"
          ]
        },
        "description": "A machine-readable registry of quantitative and qualitative claims published by Wavect, with review dates, localized page appearances and public third-party citations where available.",
        "inLanguage": "en",
        "isAccessibleForFree": true,
        "license": "https://creativecommons.org/licenses/by/4.0/",
        "name": "Wavect verified publication claims",
        "url": "https://wavect.io/verified-claims.json"
      },
      "url": "https://wavect.io/"
    },
    {
      "@id": "https://wavect.io/team/kevin-riedl/#person",
      "@type": "Person",
      "jobTitle": "Managing Director",
      "name": "Kevin Riedl",
      "sameAs": [
        "https://www.wikidata.org/wiki/Q139796365",
        "https://www.linkedin.com/in/wsdt",
        "https://github.com/wsdt"
      ],
      "url": "https://wavect.io/team/kevin-riedl/",
      "worksFor": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      }
    },
    {
      "@id": "https://wavect.io/team/christof-jori/#person",
      "@type": "Person",
      "jobTitle": "Managing Director",
      "name": "Christof Jori",
      "sameAs": [
        "https://www.wikidata.org/wiki/Q139796367",
        "https://www.linkedin.com/in/jocr77/",
        "https://github.com/jo-chris"
      ],
      "url": "https://wavect.io/team/christof-jori/",
      "worksFor": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      }
    },
    {
      "@id": "https://wavect.io/#website",
      "@type": "WebSite",
      "inLanguage": [
        "en",
        "de",
        "es",
        "zh"
      ],
      "name": "Wavect",
      "potentialAction": {
        "@type": "SearchAction",
        "query-input": "required name=search_term_string",
        "target": {
          "@type": "EntryPoint",
          "urlTemplate": "https://wavect.io/search/?q={search_term_string}"
        }
      },
      "publisher": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      },
      "url": "https://wavect.io/"
    },
    {
      "@id": "https://wavect.io/blog/arga-vs-archal-agent-integration-testing/#webpage",
      "@type": "WebPage",
      "dateModified": "2026-10-08",
      "inLanguage": "en",
      "isPartOf": {
        "@id": "https://wavect.io/#website",
        "@type": "WebSite"
      },
      "lastReviewed": "2026-10-08",
      "url": "https://wavect.io/blog/arga-vs-archal-agent-integration-testing/"
    }
  ]
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  "abstract": "Choose Arga or Archal by the exact service operations, starting state and CI lifecycle your workflow needs. Stateful twins expose cross-call effects that simple mocks can miss, but passing a simulation does not verify the real provider. Keep a separate contract smoke test.",
  "articleBody": " Blog overview/Delivery and QA/QA and production readiness Arga Labs vs Archal: Stateful Agent Integration Tests TL;DR Choose Arga or Archal by the exact service operations, starting state and CI lifecycle your workflow needs. Stateful twins expose cross-call effects that simple mocks can miss, but passing a simulation does not verify the real provider. Keep a separate contract smoke test. Evidence: Documentation reviewed on 8 October 2026. This is a researched implementation guide. The pilot below is proposed; we have not run these vendor evaluations or measured their performance. When do stateful API environments beat simple mocks? When correctness depends on what earlier actions changed. A mock returning “success” for posting a Slack message does not prove that a retry created only one message, that an unauthorized channel stayed unchanged, or that a later read saw the correct state. A stateful environment lets you inspect those consequences. The YC profile for Arga Labs identifies the company at argalabs.com. The launch describes service replicas and earlier application-staging behavior. Use current documentation to scope a purchase; do not assume a historical launch describes the current deployment contract. What is the current Arga versus Archal boundary? Arga's current test-workflow documentation separates Twin Runs from browser Test Runs and PR Test Runs. PR Test Runs use the configured application URL; they do not deploy the changed app into a per-PR Arga environment. Your CI must still supply the correct application build and reachable URL. Archal's sandbox starting-state documentation describes versioned samples and explicit JSON, with guarded SQL for Supabase. State loaded during creation establishes the reset baseline. This differs from older material describing natural-language scenario generation. Compare current state contracts, not product slogans. DecisionArga documentationArchal documentation Test your application's UIBrowser tests against a reachable URLProvide your own application and test runnerReproduce starting stateSaved scenarios and twin stateVersioned samples or validated explicit stateValidate the changed appSupply the intended deployment URLRun your app or agent against returned service endpoints What should a GitHub-to-Slack test assert? Proposed task: when an authorized issue is labeled “ready,” post one summary in the allowed Slack channel and save a durable delivery reference. Prepare the same synthetic repository, issue, channel and permission intent for both products. Confirm the required API operations are actually supported before provisioning. These are test requirements, not claims that both vendors implement every fault. CaseInjectionRequired final state Happy pathOne allowed label eventOne message and one delivery referenceDuplicate eventDeliver the same event twiceStill one logical deliveryInterrupted runStop after send, before local acknowledgementReconcile existing message before retryDenied channelRemove posting permissionNo message; explicit failureUnexpected responseAdd a controlled malformed response in the harnessNo false success; inspectable errorResetRestore the fixture and rerunSame initial records and equivalent outcome Check destination records independently of the agent's final answer. “Done” is not a state assertion. Keep an application database fixture too: resetting a vendor twin does not automatically reset your own delivery ledger. How should CI isolate credentials and failures? Use returned sandbox endpoints and scoped credentials, with an allowlist that prevents real-provider fallback. Archal's current lifecycle and authentication reference separates workspace control credentials from environment credentials. A failed environment startup should fail the job, rather than redirect test traffic to production. Save a before-state snapshot, request trace, after-state snapshot and operation ledger. In cleanup, tear down the session even when assertions fail. Reset between independent cases, and serialize tests that intentionally share state. Redact credentials and personal data before storing artifacts. What do these environments not prove? Check Arga's per-twin support and limitations for each required endpoint. Simulations can differ from providers in rate limits, OAuth, timing, webhook delivery and undocumented behavior. Maintain a narrow real-provider test account for contract checks, while using twins for repeatable fault and state tests. Passing one should not waive the other. Archal's usage documentation currently bills ready environment duration at USD 0.10 per environment-minute, prorated by seconds. Two ready environments for ten minutes model USD 2, before any other tool costs. This is arithmetic, not a measured run. Include setup, teardown, idle time and the current availability of paid continuation in your adoption discussion. Paid continuation is currently not self-serve in the checked documentation; confirm access before planning",
  "articleSection": "Engineering",
  "author": {
    "@id": "https://wavect.io/team/kevin-riedl/#person",
    "@type": "Person",
    "name": "Kevin Riedl",
    "sameAs": [
      "https://www.wikidata.org/wiki/Q139796365",
      "https://www.linkedin.com/in/wsdt",
      "https://github.com/wsdt"
    ],
    "url": "https://wavect.io/team/kevin-riedl/"
  },
  "citation": [
    {
      "@type": "WebPage",
      "name": "YC profile for Arga Labs",
      "url": "https://www.ycombinator.com/companies/arga-labs"
    },
    {
      "@type": "WebPage",
      "name": "current test-workflow documentation",
      "url": "https://docs.argalabs.com/features/validate-modes"
    },
    {
      "@type": "WebPage",
      "name": "sandbox starting-state documentation",
      "url": "https://docs.archal.ai/sandboxes/starting-state"
    },
    {
      "@type": "WebPage",
      "name": "current lifecycle and authentication reference",
      "url": "https://docs.archal.ai/llms.txt"
    },
    {
      "@type": "WebPage",
      "name": "per-twin support and limitations",
      "url": "https://docs.argalabs.com/concepts/twin-reference"
    },
    {
      "@type": "WebPage",
      "name": "usage documentation",
      "url": "https://docs.archal.ai/pricing-usage"
    }
  ],
  "dateModified": "2026-10-08",
  "datePublished": "2026-10-08",
  "description": "Choose Arga or Archal by the exact service operations, starting state and CI lifecycle your workflow needs. Stateful twins expose cross-call effects that simple mocks can miss, but passing a simulation does not verify the real provider. Keep a separate contract smoke test.",
  "headline": "Arga Labs vs Archal: Stateful Agent Integration Tests",
  "image": "https://wavect.io/img/blog/headers/header_arga-vs-archal-agent-integration-testing.svg",
  "inLanguage": "en",
  "keywords": "Engineering, AI agents",
  "mainEntityOfPage": {
    "@id": "https://wavect.io/blog/arga-vs-archal-agent-integration-testing/",
    "@type": "WebPage"
  },
  "publisher": {
    "@id": "https://wavect.io/#organization",
    "@type": [
      "Organization",
      "ProfessionalService",
      "LocalBusiness"
    ]
  },
  "url": "https://wavect.io/blog/arga-vs-archal-agent-integration-testing/",
  "wordCount": 1160
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "BreadcrumbList",
  "itemListElement": [
    {
      "@type": "ListItem",
      "item": "https://wavect.io/",
      "name": "Home",
      "position": 1
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/overview/",
      "name": "Blog overview",
      "position": 2
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/topics/delivery-qa/",
      "name": "Delivery and QA",
      "position": 3
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/clusters/qa-production/",
      "name": "QA and production readiness",
      "position": 4
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/arga-vs-archal-agent-integration-testing/",
      "name": "Arga Labs vs Archal: Stateful Agent Integration Tests",
      "position": 5
    }
  ]
}
```
