# Pricing and methodology audit

Audited 2026-09-04. This document records the evidence standard, comparison
scope, material corrections, validation strategy, and known limits for
sandboxprices.com.

`frontend/src/data.js` is the canonical provider dataset and workload.
`frontend/src/tco.js` is the canonical browser calculator. The backend dataset
is generated from the frontend data, and `backend/app/main.py` independently
implements the same formulas so parity can be tested rather than assumed.

## Trust policy

Pricing fails closed. A paid product can receive a numeric total only when it
links a current first-party pricing or billing source and every material meter
needed by the workload can be reproduced. A newly added paid row starts as
unverified. Missing prices are not replaced by averages, competitor claims, or
plausible market rates.

Evidence labels have deliberately narrow meanings:

- **Verified:** rates, billable allocation, and represented billing mechanics
  are directly supported by current first-party evidence for the displayed
  scope.
- **Modeled:** first-party rates are used, but a disclosed assumption remains.
  Typical examples are GB versus GiB semantics, an account-wide allowance,
  region selection, sampled usage, or converting aggregate monthly inputs into
  an implied average session duration.
- **Unverified:** a complete reproducible current total is not available. The
  product remains discoverable but cannot receive a numeric rank.
- **Self-hosted:** no managed-service list price applies. A meaningful total
  needs operator-selected infrastructure and operations costs.

Promotional signup credit, private invitations, temporary previews, and
negotiated contracts are excluded from gross comparisons. An account-wide
recurring credit is applied only in invoice mode and is identified as
account-wide.

## Current audit result

The catalog contains 70 products.

| Evidence state | Count | Ranking behavior |
| --- | ---: | --- |
| Verified | 1 | Numeric when the exact workload is supported |
| Modeled | 44 | Numeric when supported, with the assumption disclosed |
| Unverified | 11 | Never numeric or ranked |
| Self-hosted | 14 | No managed-service price |

The one strictly verified row is CreateOS Sandbox. The small verified set is
intentional. A provider with correct first-party rates remains modeled whenever
the monthly result still depends on an assumption or source conflict that the
public material cannot eliminate. Superserve is modeled because its live billing
endpoint marks storage non-billable while its static pricing page says paused
storage is charged.

Every one of the 16 previously unverified rows was re-researched against its
current first-party site, documentation, public API, or product catalog. Five
were false negatives and are now modeled: Sailboxes, Declaw, MIOSA, CodeSandbox
SDK, and Sandbox0. The remaining 11 have an identified material evidence gap,
listed below; they are not an undifferentiated research backlog.

For the canonical default workload:

- 37 managed products produce a numeric total.
- 5 infrastructure reference rows produce a numeric total but are separated
  from the managed-product ranking.
- AWS EC2 `t4g.nano` is the permanent raw-VM baseline. It remains visible in
  the default human table with a `ref` marker even though the other
  infrastructure references remain behind the optional self-host filter.
- Deno Sandbox, Ellipsis, and Baponi are audited but reject this default
  workload under their published session or execution constraints.
- 11 products are unverified and 14 are self-hosted.

The generated `/pricing.md` document is the current source for exact rankings,
totals, Lambda ratios, billed allocations, invoice context, and every row's
first-party link. This audit intentionally does not copy those changing values
into a second hand-maintained table.

## Canonical comparison workload

The default is anchored to AWS Lambda MicroVM's smallest published shape. This
is the comparison the site is designed to answer: how much more or less another
product costs for the smallest Lambda-sized agent box.

| Input | Default | Meaning |
| --- | ---: | --- |
| Requested CPU | 0.25 vCPU | Lambda `micro-0.25` baseline |
| Requested RAM | 0.5 GiB | Lambda `micro-0.25` baseline |
| Requested disk | 5 GiB | Fits its published 8 GB disk ceiling |
| Metered RAM | 0.5 GiB | Used only by observed-memory meters |
| Hot disk | 5 GiB | Used only by actual active-byte meters |
| Cold disk | 5 GiB | Used only by actual retained-byte meters |
| Running time | 200 hours/month | Aggregate running sandbox-hours |
| Suspended time | 50 hours/month | Explicit retained suspended state |
| Executions | 60/month | Drives per-execution minimums and rounding |
| New sandboxes | 60/month | Drives creation and base-image reads |
| Resumes | 50/month | Drives resume and state-read charges |
| Snapshot state | 2 GiB | Explicit saved-state quantity |
| Stored image | 2 GiB | Explicit base-image quantity |
| Actual CPU utilization | 50 percent | Used only by actual-CPU meters |
| CPU and RAM burst | 1x | Baseline use for baseline/peak products |

Five GiB is intentional. A 10 GiB default would force Lambda to `micro-2`
because `micro-0.25`, `micro-0.5`, and `micro-1` each have an 8 GB disk ceiling.
Other providers receive the same 0.25 vCPU, 0.5 GiB, 5 GiB request and either
accept it, round it up to the smallest valid single-sandbox allocation, or
return unsupported. Multiple small boxes never substitute for one requested
box.

## Cost rules

- Active hours always mean running sandbox-hours. They never become 730 hours
  merely because idle compute is billable. Only rows explicitly marked as
  always-on raw VM references use the full 730-hour month.
- CPU utilization affects only providers that meter actual CPU consumption.
  Allocated CPU, active-runtime, fixed-box, and RAM-second products use their
  published basis.
- Memory is priced independently unless a provider publishes a fixed shape or
  a CPU-to-memory rule.
- Resource constraints resolve before pricing. Requests round upward to valid
  values, and unsupported shapes return no total.
- Provisioned disk and actual stored bytes are different inputs. A retained
  allocation bills for the full month only when the provider publishes that
  lifecycle behavior. Included ephemeral disk never acquires an invented
  storage rate.
- Snapshot state and stored base images are separate. Snapshot writes, state
  reads, image retention, and launch reads use only the meter to which the
  provider assigns them.
- Run, session, suspend, and lifecycle counts are separate. Aggregate hours are
  divided evenly only where a per-event minimum or maximum makes an implied
  average necessary, and that assumption forces a modeled label.
- Plans must satisfy resource, session, and usage eligibility. Independent hard
  caps remain independent; spare memory allowance cannot conceal excess CPU.
- Sub-cent components remain unrounded internally. Display formatting occurs
  only after the full total is calculated.
- Decimal GB and binary GiB are converted by bytes when a provider defines its
  meter that way. An undefined provider convention is stated as a 1:1 modeling
  assumption rather than silently converted.

## Gross comparison and invoice context

The primary ranking and every `vs Lambda` ratio use gross usage for the exact
same workload. Gross usage is the provider's public resource meter before
optional plan fees, account-wide recurring credits, allowances, and monthly
minimums. An inseparable fixed plan or bundle remains in gross usage because no
separate resource-only rate exists.

Invoice mode applies the cheapest eligible public plan that the calculator can
reproduce, including required fees, recurring allowances, hard caps, and usage
minimums. It is useful purchase context, but it is not the comparison ranking:
an account-wide allowance may already be consumed by another workload.

The machine-readable pricing document publishes both `gross_usd_month` and
`modeled_invoice_usd_month`, and names the selected invoice context. This keeps
agents from mistaking a marginal resource subtotal for the minimum bill needed
to use products such as exe.dev.

## AWS Lambda MicroVM reference calculation

The default selects `micro-0.25` without rounding. The current US East
(N. Virginia), Graviton model uses the published baseline rates and separate
state meters:

- CPU: `$0.09969984 x 0.25 vCPU x 200 h = $4.984992`.
- Memory: `$0.01320012 x 0.5 GB x 200 h = $1.320012`.
- Stored 2 GB base image: `$0.08 x 2 = $0.16`.
- Suspended 2 GB state: `$0.08 x 2 x 50 / 720 = $0.011111111`.
- 50 suspend/resume cycles: `($0.0038 + $0.00155) x 2 x 50 = $0.535`.
- 60 new launches reading a 2 GB image: `$0.00155 x 2 x 60 = $0.186`.
- Gross default total: `$7.197115111` per month.

CPU and RAM burst are modeled independently and capped at the published 4x
peak. A MicroVM lifecycle may span at most eight hours across running and
suspended states, and an individual suspension is also limited to eight hours.
The calculator checks implied average durations from the aggregate inputs;
because it cannot prove every individual duration, this row remains modeled.
Data transfer is excluded.

## Material corrections retained by regression tests

- Rebuilt Lambda from its published per-second CPU and RAM rates, independent
  CPU and RAM burst, disk-dependent shapes, image retention, suspended-state
  prorating, snapshot writes and reads, launch reads, and lifecycle limits.
- Changed the default from 1 vCPU, 2 GiB, 10 GiB to Lambda's actual smallest
  0.25 vCPU, 0.5 GiB, 5 GiB comparison request.
- Removed every synthetic fallback component rate. Missing a material price now
  un-ranks the row.
- Separated gross usage from invoice context so required plans, recurring
  credits, independent quotas, and monthly minimums cannot leak into the
  provider meter comparison.
- Added exact resource resolution for fixed shapes, CPU shapes, resource steps,
  CPU-to-memory ratios, capacity-only validation, provider-specific decimal-GB
  RAM and storage, and unsupported requests.
- Replaced Northflank's impossible independent CPU/RAM request with its actual
  deployment-plan catalog. The default now selects `nf-compute-50` and bills
  its 0.5 vCPU / 1024 MB allocation.
- Corrected Upstash's box ceilings and storage meter to literal decimal GB. A
  5 GiB request is 5.36870912 GB, so the default selects Medium rather than
  incorrectly fitting Small's 5 GB ceiling.
- Corrected actual versus provisioned CPU, RAM, hot disk, cold disk, snapshot,
  and image billing so a slider affects only providers that publish that meter.
- Corrected account-wide storage allowances for Novita, Deno, Archil, Runloop,
  and other plan products so they apply only in invoice mode and never erase
  gross usage.
- Rebuilt Vercel plan eligibility and independent Hobby quotas, per-session RAM
  minimums, creations, and snapshot storage.
- Rebuilt Cloudflare custom and predefined shape selection with byte-exact disk
  conversion while excluding Workers and Durable Objects costs that have no
  matching workload input.
- Rebuilt Fly Machines from exact Ashburn per-second preset rates, integer
  decimal-GB volumes, stopped rootfs, and snapshot allowances.
- Rebuilt Fly.io Sprites from sampled active CPU, memory, hot blocks, and cold
  blocks using explicit observed-use inputs, byte-exact decimal-GB billing, its
  8-vCPU execution capacity, and literal 100 GB storage ceiling.
- Rebuilt Cloud Run as a Jobs reference with its compatible CPU/RAM shapes,
  10 GiB provisioned ephemeral-disk minimum, free tier, and one-minute
  per-task minimum under an equal-duration assumption.
- Rebuilt Morph's MCU maximum-dimension formula and snapshot capacity billing,
  Box by ASCII's whole-VM tiers and account minimum, and Baponi's per-execution
  credit rounding.
- Corrected LangSmith's LCU/LSU conversion, Modal's CPU and memory rates,
  Railway's per-minute VM beta rates, Tenki's active and retained disk meters,
  and GKE Autopilot's CPU-to-memory constraints.
- Corrected Daytona's public limits to 4 vCPU, 8 GiB RAM, and 10 GiB disk.
- Rebuilt Deno around its fixed 2-vCPU capacity, 768 MB memory floor,
  configured-memory billing, active-CPU billing, current plan allowances, and
  strict 30-minute session ceiling.
- Rebuilt Runloop with its current CPU, RAM, and storage rates, discrete custom
  resource steps, 1:2 through 1:8 CPU-to-memory ratio, 48-hour Devbox lifetime,
  and Pro requirement for suspend and resume.
- Corrected Blaxel to its RAM-derived CPU allocation and separated provisioned
  volume, standby snapshot, and stored image meters. The requested disk is a
  volume because the half-RAM writable root tmpfs cannot satisfy it.
- Added AgentCore's published 128 MB measured-memory billing floor and corrected
  its isolation description to dedicated microVM without inventing a VMM.
- Corrected Fly Machines snapshot treatment: automatic volume snapshots are
  optional and are disabled in the modeled configuration because incremental
  snapshot bytes are not a workload input.
- Removed a nonexistent `s-2vcpu-4gb` CreateOS shape and removed an unpublished
  startup claim. Only its four current first-party shapes can be selected.
- Changed Archil's provider domain to `archil.com` and its isolation type to
  unknown. Archil publishes persistent sandbox behavior and pricing but does
  not currently disclose the VMM.
- Restored Sailboxes using its complete first-party observed CPU, RAM, and NVMe
  disk meters, the S/M/L capacity table, and each size's distinct creation fee.
  The default selects S but bills observed usage rather than its ceilings.
- Restored Declaw using its template resource ranges, fixed 20 GB overlay,
  provisioned-resource meters, and plan limits. The fixed disk requires Pro in
  invoice mode, whose $100 monthly recharge is a carried-balance commitment.
- Restored MIOSA from the exact credit-denominated component rates and its
  unauthenticated live Sandbox shape contracts. The default selects Tiny and
  charges its 20 GB disk at the running and cold rates across the month.
- Restored CodeSandbox SDK for the current Nano shape only. Its first-party
  pricing page directly publishes the $0.15 on-demand VM-hour price and 40/160
  Nano-equivalent included hours; unpublished larger-tier multipliers remain
  outside the model.
- Restored Sandbox0 with its directly documented 500m CPU / 512Mi memory / 8Gi
  custom-template configuration, configured-memory meter, and actual retained
  rootfs byte-time meter. This configuration is not presented as the provider's
  minimum.
- Kept Run Cloud unranked because storage is explicitly metered but no public
  storage rate is listed.
- Unranked Railway Sandboxes because the VM beta rates are public but the
  sandbox SDK and product documentation do not publish CPU or memory capacity,
  a billing floor, or enough fields to validate the Lambda-sized allocation.

## Deliberately unranked products

| Product | Why no numeric total is shown |
| --- | --- |
| Scrapybara | The former product is superseded by Capy and has no current service rate |
| Railway Sandboxes | VM beta rates are public, but sandbox capacity and the billing floor are not |
| Computer Agents | Durable retained-storage capacity or overage is not reproducible |
| Arker | Invite-only paid access has no public rate card |
| Mosaic | The vendor is identified, but it publishes no public rate card |
| Run Cloud | Storage is metered but its public storage rate is missing |
| Sandbox as a Service | Current plans omit the included-hours and usage relationship needed for TCO |
| Leap0 | Preview access is not a durable paid rate |
| OmniRun | Disk fit, retained-state equivalence, and a time-stamped USD conversion are missing |
| DeepInfra Sandboxes | Exact catalog shapes require authenticated API access |
| CoreWeave Sandboxes | Pricing depends on contracted cluster capacity rather than a public sandbox rate |

## Reproducible verification

The repository enforces the following on every pricing change:

- Backend data must regenerate exactly from the canonical frontend dataset.
- Frontend and backend golden tests cover billing components, lifecycle rules,
  shapes, plans, allowances, unsupported workloads, and fail-closed rows.
- `scripts/verify-invariants.mjs` checks evidence gates, source quality,
  nonnegative finite totals, shape validity, billed allocations, monotonicity,
  generated discovery-file freshness, canonical metadata, and Dataset JSON-LD.
- The invariant and UI suites explicitly require the `ec2-t4g-nano` row to
  exist, price the canonical workload, retain its always-on reference status,
  and remain visible by default.
- `scripts/verify-parity.mjs` exercises representative and boundary workloads
  across every product in gross and invoice modes, comparing statuses, totals,
  allocations, components, allowances, lifecycle charges, and selected plans.
- The production frontend build must complete after the tests.
- Source and generated files are scanned for forbidden em dash characters.

The required commands are documented in `AGENTS.md` and run in pull-request and
release workflows. A pricing change is not considered complete until all of
them pass.

Last full local verification on 2026-09-04: 73 frontend tests, 67 backend
tests, 691,040 frontend/backend product-workload comparisons with no
differences, a successful production build, successful invariant and generated
file checks, valid sitemap XML, and a successful `nginx -t` against the
production configuration.

## Agent and crawler publication

The human calculator is accompanied by generated, source-linked machine
resources:

- `/pricing.md` is the canonical agent-readable default comparison.
- `/pricing.txt` is a byte-identical compatibility alias that identifies
  `/pricing.md` as canonical.
- `/llms.txt` is a concise agent discovery index. It is an emerging convention,
  not a crawl-permission standard.
- `/robots.txt` allows all compliant crawlers, explicitly repeats the current
  OpenAI, Anthropic, Perplexity, and Google AI crawler controls, and points to
  the sitemap. The wildcard remains authoritative for new or unnamed agents.
- `/sitemap.xml` lists only canonical public documents.
- The homepage declares its canonical URL, Markdown alternate, `llms.txt`
  description, and Schema.org `Dataset` JSON-LD distributions.
- The API subdomain independently serves `/robots.txt`, and its root document
  links the pricing, methodology, product, TCO, and OpenAPI resources.

Robots rules are advisory and cannot force indexing. After deployment, the CDN,
WAF, and origin must also return successful responses to crawler user agents.
Search-engine submission and URL inspection remain operational deployment
steps, not source-code metadata.

## Known limits

No static calculator can guarantee a future invoice for every account. This
site aims for technical validity within an explicit scope:

- Prices can change after the audit date and need recurring first-party review.
- Region, currency, taxes, egress, requests, control-plane charges, concurrency,
  model tokens, negotiated discounts, and surrounding cluster costs are outside
  a row unless its explanation includes them.
- Account-wide allowances may already be consumed elsewhere. Invoice mode
  assumes the displayed recurring allowance is available to this workload.
- Aggregate time cannot prove each individual run, session, suspension, or
  lifecycle duration. Any row depending on an implied average is modeled.
- Sampled or peak consumption meters can differ from a monthly average when the
  real time series is bursty.
- EC2 T4g surplus-credit behavior is evaluated from a monthly-average workload,
  not AWS's rolling 24-hour credit timeline.
- One retained disk, snapshot, and base image are modeled. Fleets with multiple
  retained copies must represent those bytes separately.
- A first-party page can still be ambiguous or internally inconsistent. The row
  is modeled or unranked until the ambiguity is resolved.

Corrections are welcome at <hello@sandboxprices.com>. A useful report includes
the provider, exact first-party URL, billing region and currency, and the rate
or formula that changed.
