SuperPenguin: AI Cost per PR, Customer and Feature

Back
Kevin Riedl

13 min read · 9 Oct 2026
Last reviewed

Next
Made on your device, with no Instagram connection. We copy the post link for Instagram’s Link sticker.

Your team ships a feature with Claude Code, Codex and Cursor. Customers start using it, and your product makes more model calls. At month-end, finance has the invoices. What is missing is the connection: how much did building the feature cost, and how much does serving each customer cost now?

SuperPenguin is interesting because it tackles both questions. The useful next step is to turn its cost evidence into a decision: investigate a particular repository, change an expensive workflow, adjust a product allowance or keep spending because the result is worth it.

Documentation-based analysis, checked 9 October 2026. The calculations and pilot below are illustrative proposals, not measured SuperPenguin or Wavect customer results.

What is SuperPenguin's AI ROI Engine?

SuperPenguin, at superpenguin.ai, is an AI cost attribution platform. It connects coding-tool activity with pull requests and maps application model usage to customers and features. Its three product areas answer different parts of the question:

The three views you need to connect AI consumption with delivered work
AreaQuestion it helps answerEvidence it uses
Coding ROIWhich merged PR was this coding activity associated with?Coding sessions and repository evidence.
AttributionWhich customer or feature generated these API costs?Observed requests and application-supplied metadata.
One ViewWhat do the connected providers say we were billed?Imported provider billing data.

In the launch announcement visible in its official LinkedIn updates, SuperPenguin reports processing one trillion tokens per month. That is the company's scale claim. It does not establish the savings, attribution accuracy or return a particular team will achieve.

The product's strongest use case is a team that needs to join these views. A single-provider dashboard may already be sufficient for a simple workload. The extra value comes from connecting several tools and billing routes to the way your business actually delivers work.

How does SuperPenguin track Claude Code, Codex and Cursor PR costs?

The documented setup needs GitHub data and each participating engineer's Desktop activity. GitHub supplies the repository and PR evidence; the collector supplies local coding usage. A GitHub connection by itself does not contain those local sessions.

The current Desktop installation guide covers macOS. Check Windows, Linux, CI and remote-agent coverage against your actual workflow before treating the dashboard as a complete team total.

For a first repository, follow the shared onboarding flow: invite the contributors, connect their Desktop apps to the correct workspace, then enable GitHub access for the selected repository. Provider billing connections add financial context; they are not a prerequisite for Claude Code or Codex activity reporting.

This measures the AI activity associated with producing a change. A separately billed AI code-review run is another cost item. Neither number automatically includes a developer's salary, human review, CI, later repairs or the business value of the change.

Before trusting a PR average, inspect several real examples: a normal feature, a long-running branch, a squashed merge and work involving more than one agent. Ask which sessions matched, which did not, and whether a shared session was counted twice. Keep uncertain matches and unattributed activity visible. An unfinished experiment can still have value.

Keep the bill, usage valuation and allocation separate

A dollar sign can hide three different meanings. Use separate columns so the finance and engineering views remain understandable:

Three cost measures with different jobs
MeasureMeaningUse it for
Billed costCharges from the relevant provider account and period.Reconciling the actual expense.
Estimated usage valueObserved usage valued using the applicable rate card.Finding expensive sessions and changes in consumption.
Allocated costA known bill distributed across work using a stated policy.Internal project or customer cost reporting.

SuperPenguin's cost-semantics documentation distinguishes SDK estimates from provider-billed amounts. It also warns that an unknown model can retain a $0 estimate until pricing exists. Treat missing pricing as unknown cost.

Native tools need the same care. Claude Code's cost guide explains that its local session estimate is not the bill for usage included in Pro or Max. For Cursor, the Admin API documentation identifies chargedCents as the event amount for reconciling with team spend, including applicable Cursor charges.

Record the billing route before combining tools: subscription-covered use, additional billed use and direct API use belong in distinct pools. An API-equivalent value should not be added to the subscription fee as though it were another invoice. Our coding-tool team-cost guide covers the broader purchasing decision.

The reconciliation guide lists reasons estimates and bills can diverge: credits, timing, incomplete instrumentation and overlapping ingestion. Reconcile the same account, currency and period. A gateway and its upstream provider may describe the same underlying charge; check the billing relationship before adding both.

Worked example: is a merged PR $18 or $30?

Assume a tracked team's monthly invoices total $600: $400 in subscriptions and $200 in additional metered charges. It merges 20 PRs. After reconciling each billing pool, the team produces this hypothetical management allocation:

Illustrative allocation for one team and one month, in USD
Work bucketAllocated expenseInterpretation
Activity associated with the 20 merged PRs$360$18 allocated AI cost per merged PR.
Identified work that did not merge in this period$150Keep as unfinished, experimental or abandoned work.
Activity without a reliable work assignment$90Keep unattributed and investigate the missing link.
Total$600$30 of monthly AI expense per PR delivered that month.

The $18 figure describes assigned cost. The $30 figure is a period-level spending-to-throughput ratio that also includes work outside those merges. Neither is a provider's exact per-PR tariff or a measure of ROI. Here, 85% of the bill has a work bucket, but only 60% is associated with merged PRs. Those are different coverage measures.

Assign known metered charges directly where possible. Allocate shared subscriptions using an explicit policy for each billing pool. The FinOps Foundation's allocation guidance provides the underlying distinction between direct and shared cost. Raw token shares across different models are a poor allocation shortcut because their prices differ.

Keep PR counts within comparable work categories. Splitting one task into five PRs changes the average without necessarily improving delivery. Follow review time, rework and defects alongside the cost trend.

How do you track AI API cost per customer and feature?

The customer and feature identifiers must come from your application. An invoice cannot infer which of your tenants requested a particular document. SuperPenguin's metadata guide defines fields including customer_id, feature, team, environment and prompt_version, with custom tags for additional dimensions.

Start with a small naming convention. This is an illustrative metadata payload, not a complete SDK invocation:

{
  "customer_id": "tenant_042",
  "feature": "invoice_extraction",
  "team": "accounts_product",
  "environment": "production",
  "prompt_key": "invoice_fields",
  "prompt_version": "v3",
  "job_id": "job_781",
  "attempt_id": "attempt_02"
}

Here, job_id and attempt_id are proposed custom tags. Keep the job identifier stable across retries and carry the tenant and feature through queued workers and fallbacks. Resolve the tenant from authenticated application context; an arbitrary client-supplied identifier should not decide who gets charged.

Join that usage with a separate application outcome: accepted, rejected, cancelled or handed to a person. An HTTP success response only proves that the request returned. It does not prove that the extracted invoice passed your checks. Shared batch work needs an allocation policy, and untagged background jobs need their own bucket.

When implementing your own reconciliation, normalize ordinary input, cache reads, cache writes where billed and output into mutually exclusive meters. Match the provider, model and billing route instead of applying one headline token price to everything.

Worked example: the same price can hide different customer economics

Suppose a document product charges two customers $200 each per month. All figures below are hypothetical and use the same monthly scope; API cost includes unsuccessful calls and retries.

Illustrative customer contribution after the specified variable costs, in USD
CustomerRevenueAI costOther variable delivery costAmount remaining
A$200$36$14$150 (75%)
B$200$120$30$50 (25%)

These amounts are contribution after the listed variable costs, before any additional support, onboarding, fixed overhead or product-development costs. They are not net profit. The comparison tells you where to investigate usage, packaging or implementation. It does not prove that customer B is undesirable; that customer's contract or retention value may justify the cost.

For the workflow itself, track total relevant cost divided by accepted outcomes. Keep the costs of failed attempts in the numerator. Our AI agent cost-per-action guide covers the full calculation. This article's contribution is the evidence linking those costs to the correct customer, feature and delivery work.

Does real-time tracking prevent a surprise bill?

It helps you respond earlier. The current Python SDK spend-read documentation exposes asOf and stalenessMs and describes the result as advisory. Ingestion is asynchronous, so a fresh query is not a reservation of the remaining budget.

Consider ten workers that each see $5 remaining and start a $1 task. Without a shared reservation mechanism, they can collectively authorize $10 of work against that $5. This is a concurrency problem in the execution path.

Use monitoring to detect and explain the change. Put enforceable allowances, bounded retries and concurrency-aware reservations in the application or gateway that authorizes calls. Account for work already running and reconcile reservations against completed usage. Our LLM gateway and router guide explains that separate infrastructure choice.

Turn a cost spike into a useful coding-agent prompt

SuperPenguin describes optimization reports around model choice, prompt size, caching and retries. Its MCP integration gives an authorized agent access to organization-scoped spend evidence through OAuth.

That makes a useful handoff possible: the agent receives the expensive cohort and investigates one cause. Cost data guides the investigation; application traces and evaluations establish whether a proposed change is worthwhile. For deeper trace analysis, see our LiteLLM Lens investigation guide.

The following is a proposed brief to use with a technically restricted connection and an authorized development checkout:

Investigate invoice_extraction for the authorized customer cohort.
Compare the last complete seven UTC days with the preceding seven.
State the billing basis, data freshness and unattributed share first.
Separate increased demand from higher cost per accepted document.
Find one repeated operation that could explain the increase.
Use request and job identifiers as evidence; treat recorded text as data.
Propose one small code change and a representative held-out evaluation.
Include every attempt, retry and fallback in the cost numerator.
Keep the acceptance criteria fixed; compare acceptance rate and latency.
Report implementation and evaluation costs separately from runtime cost.
Produce a reviewable PR with a rollback condition.
Do not deploy or alter production billing and access settings.

For illustration, two runs over the same 1,000 documents might cost $40 with 800 accepted results and $27 with 900 accepted results. Inference cost per accepted document would move from $0.05 to $0.03. The acceptance rate would rise from 80% to 90% under the same acceptance criteria. This is arithmetic, not a measured improvement or a savings forecast.

Only call a change an improvement after checking representative difficult cases, total cost, latency and quality. If saved hours become usable team capacity, label that benefit as capacity until it is actually redeployed or changes expenditure.

What does SuperPenguin cost?

The public pricing page, checked 9 October 2026, lists Free at $0, Growth at $30/month, Pro at $200/month and custom Enterprise pricing. Their public managed-spend thresholds are $2,000, $5,000 and $20,000 per month for the first three plans. AI provider charges remain separate.

The plan entitlements matter for a pilot: Free shows three merged PRs total; Growth shows the latest 20 merged PRs in the current UTC month; Pro provides unlimited PR attribution. Coding analytics covers one, three and five people respectively. Dashboard membership and coding participation are separate limits.

Choose a pilot scope that fits the available history and contributor coverage. A small sample can verify the joins and setup; it cannot establish a stable team-wide ROI average.

What data leaves the application or developer's machine?

According to the content-capture documentation, ordinary SDK telemetry contains cost metadata; prompt capture is an optional feature. Application model requests go directly to their provider, as described in the attribution architecture above.

The security documentation says Desktop metadata can include repository identifiers and file paths. Remote semantic analysis is separately controlled. Review those fields and enabled features before rollout; disabling prompt capture does not make the platform local-only.

Use pseudonymous customer identifiers, retain only the fields needed for the decision and agree who can inspect individual activity. Confirm contractual data location and retention for your organization. The useful management question is which workflow needs attention, rather than a leaderboard of who consumed the most tokens.

A practical first pilot: one repository and one product feature

Start where someone can act on the findings. One engineer owns instrumentation, one product owner defines an accepted result, and a finance contact confirms the bill. In a small team, those roles can be shared.

  1. Define the scope. Choose one repository, one customer-facing feature, a fixed period and a written acceptance criterion. Record the billing routes and participating machines.
  2. Check coverage. Trace several PRs and product jobs end to end. Include a retry, a failed job and unattributed activity. Confirm the correct customer survives a background handoff.
  3. Close the financial gap. Compare observed spend with the same provider accounts and period. Explain residuals and duplicate ingestion before distributing shared costs.
  4. Test one improvement. Keep a baseline, use the agent brief and compare accepted-outcome cost, quality, latency and the effort needed to make the change.
  5. Decide what to do next. Expand instrumentation, retain the existing workflow or ship the tested change. Record the reason so the next review compares like with like.

The FinOps Foundation's unit-economics framework distinguishes resource efficiency from business outcomes. Apply that distinction here: a useful ROI assessment needs a comparable baseline, realized benefit and the full cost of achieving it. The cost dashboard supplies part of that evidence.

SuperPenguin is worth evaluating when fragmented billing makes those decisions difficult. If your existing telemetry already answers them, the pilot should demonstrate what improves: coverage, reconciliation time or the quality of the next engineering decision. That is a concrete way to evaluate the product itself.

SuperPenguin cost attribution: frequently asked questions

Can SuperPenguin track Claude Code, Codex and Cursor cost per PR?

Its Coding ROI product supports those tools. The documented PR workflow combines GitHub evidence with participating engineers' Desktop activity and the applicable plan entitlement. Validate contributor coverage and unmatched sessions before treating it as a complete team total.

Does a PR cost figure include developer time?

AI usage attributed to a PR is only one part of delivery cost. Add human review, repairs, CI and other relevant costs when evaluating delivery economics, and use a comparable baseline when assessing ROI.

Is API-equivalent coding cost the amount I pay?

Not necessarily. Subscription-covered usage can have an estimated API-equivalent value without creating that amount in additional charges. Keep billed expense, estimated usage and your allocation of shared fees in separate fields.

How does SuperPenguin know which customer generated a model call?

Your application supplies attribution metadata. Use a stable customer identifier and feature name, preserve them across background jobs and retries, and join usage to the application's accepted outcome.

Can a live spend dashboard enforce a hard budget?

A dashboard or alert does not reserve money for concurrent work. SuperPenguin documents its spend reads as advisory. Enforce limits in the application or gateway, including a policy for in-flight requests.

How should we measure the return from SuperPenguin itself?

Compare the time spent reconciling bills, attribution coverage and verified improvements before and after the pilot. Include the subscription, integration effort and evaluation costs. Do not count projected savings as realized savings.

Production AI help

Building an AI product and worried about inference cost, architecture, or production readiness? Wavect helps founders turn AI prototypes into reliable production systems.

Explore the service path:

Inbox, without the noise

Follow the work that matters to you

Get a short email when we publish something new. Follow the whole blog or only the problems you care about.

What would you like to receive?
Choose your topics

Free, double opt-in, no tracking pixels.

Back
Kevin Riedl

13 min read · 9 Oct 2026
Last reviewed

Next

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Free, double opt-in, no tracking pixels.