Simulation Accuracy Benchmark
This page publishes a reproducible, company-owned cwm-bench measurement of the canonical AWS architecture and clearly separates it from engine predictions and provider-documentation references.
See also: Simulation Fidelity — benchmark data and accuracy ranges for all five cloud providers.
AWS scored app CPU, in-VPC internal-load-balancer latency, goodput and CRUD errors come from the pinned owned cwm-bench campaign. Cost uses the us-east-2 price list. Tuning uses 10 / 100 / 500 RPS; 1,000 RPS, later-day and us-west-2 are holdouts. See the Methodology section below for full source citations.
GCP, Azure, OCI, and DigitalOcean scores use documentation references, not owned measurements. They are shown separately from canonical AWS and are not a cross-provider ranking.
Independently sourced coverage currently includes rosa-hcp, rosa-classic, aro, openshift-dedicated, self-managed across 6 region-specific scenarios. Platform fees are estimated from published product/pricing information; performance behavior is not a topology-matched public load test.
View source-backed OpenShift reference output| Provider | Scoring basis | Overall | ||||
|---|---|---|---|---|---|---|
| Owned lean-weight measurement comparison (cost: public price list); typical holdout shown separately below | ||||||
AWS (lean) | Measured (CWM Bench)Measured performance and goodput/error; cost is a public price-list reference, not a measured bill. | 93.8% | ||||
| Documentation-reference comparisons — not yet measured | ||||||
GCP | Documentation reference (not yet measured)Documentation references have not been checked against owned measurements. | 98.0% | ||||
Azure | Documentation reference (not yet measured)Documentation references have not been checked against owned measurements. | 97.8% | ||||
Oracle Cloud (OCI) | Documentation reference (not yet measured)Documentation references have not been checked against owned measurements. | 96.2% | ||||
DigitalOcean | Documentation reference (not yet measured)Documentation references have not been checked against owned measurements. | 97.4% | ||||
Click any row to load that provider's full benchmark results below.
| Provider / scoring basis | Idle 10 req/s | Normal 100 req/s | Peak 500 req/s | Burst 1,000 req/s |
|---|---|---|---|---|
| Owned measurement comparison | ||||
| AWSMeasured (CWM Bench) | 0.4%measured app CPU 0.5% | 1.9%measured app CPU 1.8% | 8.6%measured app CPU 8.6% | 17.0%measured app CPU 13.4% |
| Documentation references — not checked against owned measurements | ||||
| GCPDocumentation reference (not yet measured) | 6.3%doc ref 6.0% | 22.0%doc ref 22.0% | 57.3%doc ref 58.0% | 88.1%doc ref 84.0% |
| AzureDocumentation reference (not yet measured) | 5.4%doc ref 5.0% | 18.5%doc ref 19.0% | 52.3%doc ref 54.0% | 86.1%doc ref 81.0% |
| Oracle Cloud (OCI)Documentation reference (not yet measured) | 4.3%doc ref 4.0% | 12.1%doc ref 12.0% | 44.4%doc ref 48.0% | 82.6%doc ref 78.0% |
| DigitalOceanDocumentation reference (not yet measured) | 6.3%doc ref 6.0% | 15.4%doc ref 15.0% | 41.9%doc ref 44.0% | 70.3%doc ref 68.0% |
AWS cells show predicted application-server CPU against measured application-server CPU; other providers show simulated output against documentation-sourced references. Click any row to load that provider's full benchmark below.
Select a provider to load its canonical benchmark automatically. Measured (CWM Bench) — Measured performance and goodput/error; cost is a public price-list reference, not a measured bill.
Lean Overall Simulation Accuracy: 93.8%
Weighted composite across P50/P95 latency, CPU utilization, throughput, error rate, and cost — averaged over four traffic scenarios (Idle, Normal, Peak, Burst).
Lean burst holdout (not used for tuning): 85.4% · 3 tuning runs + 1 holdout
This 93.8% overall result and burst holdout measure lean app weight. Self-serve and the GitHub App default to typical; its separate owned 300 RPS holdout below is not included in the lean overall score.
Throughput largely follows the prescribed workload; cost uses the AWS public price list, not a measured bill.
Measured performance and goodput/error; cost is a public price-list reference, not a measured bill.
Idle
95%
Normal
95%
Peak
100%
Burst
85%
300 RPS holdout — before fit: 24.3/100 (27.0 without cost)
300 RPS independent holdout — after fit: 84.2/100 (82.4 without cost)
Engine 1.2.6; aws-crud/typical:typical-v1-20260927c:6aa574d7ff9d3080b88b221bcd59f7d218ae37f0. Only typical-fit-20, typical-fit-100, typical-fit-200 were used for fitting. Neither 300 nor 500 RPS was used to tune the fit.
| Metric | Predicted | Measured / reference | Metric score | Weight | Weighted contribution |
|---|---|---|---|---|---|
| P50 | 4.17 ms | 4.62 ms | 90.1 | 0.20 | 18.0 |
| P95 | 12.78 ms | 26.92 ms | 47.5 | 0.25 | 11.9 |
| App CPU | 14.74 % | 15.30 % | 96.3 | 0.20 | 19.3 |
| Throughput | 262.60 RPS | 262.61 RPS | 100.0 | 0.15 | 15.0 |
| Error rate | 0.0000 % | 0.0044 % | 100.0 | 0.10 | 10.0 |
| Cost | 0.454581 USD/hour | 0.4545 USD/hour (reference) | 100.0 | 0.10 | 10.0 |
Throughput: whole-run goodput (5 min warmup plus 15 min steady); the prediction is a point rate.
Cost: Cost was not measured. USD 0.4545/hour is a public list-price reference, not a bill. The post-fit scored prediction excludes modeled egress; the historical before-fit prediction included it, so their cost scores are not on identical bases.
500 RPS — diagnostic, extrapolated beyond fitted range, not in the headline: 71.0/100 (67.9 without cost)
Engine 1.2.6; aws-crud/typical:typical-v1-20260927c:6aa574d7ff9d3080b88b221bcd59f7d218ae37f0. Only typical-fit-20, typical-fit-100, typical-fit-200 were used for fitting. This point was not used to tune the fit; the 300 RPS result remains an independent holdout.
| Metric | Predicted | Measured / reference | Metric score |
|---|---|---|---|
| P50 | 3.451 ms | 5.188 ms | 66.5 |
| P95 | 16.333 ms | 88.190 ms | 18.5 |
| App CPU | 24.264 % | 26.342 % | 92.1 |
| Throughput | 437.588 RPS | 437.577 RPS | 100.0 |
| Error rate | 0.000 % | 0.010 % | 97.2 |
| Cost | 0.451107 USD/hour | 0.4545 USD/hour (reference) | 99.3 |
Throughput: whole-run goodput (5 min warmup plus 15 min steady); the prediction is a point rate. Cost: USD 0.4545/hour is a no-egress public list-price reference, not a measured bill; scored prediction excludes modeled egress.
One campaign, one day, us-east-2, one attempt per rung; engine 1.2.3 predictions frozen before measurement (seed 20240601, step 6). Measured cost missing. Internal ALB + 2 × m5.large + db.r5.large MySQL 8.0; two workers and pool 250 per server.
| Metric | Predicted | Measured / reference | Metric score | Weight | Weighted contribution |
|---|---|---|---|---|---|
| P50 | 17 ms | 4.62 ms | 0.0 | 0.20 | 0.0 |
| P95 | 42 ms | 26.92 ms | 44.0 | 0.25 | 11.0 |
| App CPU | 59.553 % | 15.30 % | 0.0 | 0.20 | 0.0 |
| Throughput | 299 RPS | 262.61 RPS | 86.1 | 0.15 | 12.9 |
| Error rate | 0.284 % | 0.0044 % | 3.6 | 0.10 | 0.4 |
| Cost | 0.94 USD/hour | 0.4545 USD/hour (reference) | 0.0 | 0.10 | 0.0 |
P99 (unweighted context, not part of the score): 236 vs 97.5 ms; score 0.
Throughput: whole-run goodput (5 min warmup plus 15 min steady); the prediction is a point rate.
Cost: Cost was not measured. Score uses the list-price reference of USD 0.4545 per hour (ALB 0.0225 + 2 × 0.096 + 0.24, no egress, no generator). Reference, not a bill.
Optional, unscored steady-window context: 300.0 RPS gives throughput 99.7 and total 26.3.
Historical pre-fit context — fit data, not independent holdout accuracy; old saturation diagnostic is not the extrapolated score above
- fit-20 (fit data, not holdout): 23.0 (15.3 without cost)
- fit-100 (fit data, not holdout): 19.8 (14.9 without cost)
- fit-200 (fit data, not holdout): 16.1 (14.8 without cost)
- saturation-500 (diagnostic only): 25.6 (28.5 without cost)
Sources: cwm-bench typical report · summary export · scores · frozen predictions · measurement commit 6aa574d7ff9d
Results update the architecture diagram and scenario tables below.
Canonical Architecture
| Role | Instance Type | Count | Cost / hr | |
|---|---|---|---|---|
Load Balancer | ALB (Application Load Balancer) | 1 | $0.0225 | |
Web / App Server | m5.large | 2 | $0.1920 | |
Database | db.r5.large (MySQL 8.0, Single-AZ) | 1 | $0.2400 | |
| Total (official AWS pricing, us-east-2) | $0.4545/hr | |||
Want to run this exact scenario yourself? Open the Simulation Workspace and configure the same resources.
Traffic Scenario Comparisons
Each table shows Reference vs. Simulated values, the percentage delta, and an accuracy badge (green ≥ 90%, yellow ≥ 75%, red below 75%).
Error-rate accuracy: absolute simulated-versus-reference difference ≤ 0.01 percentage points scores 100%. Above that, the score is the higher of 100 × 0.01 ÷ difference (in percentage points) and the usual relative-accuracy score (100 minus absolute percent difference, floored at 0; zero if the reference is zero and the prediction is not).
Tuning runs (fit)
| Metric | Reference scored source | Simulated | Delta | Accuracy |
|---|---|---|---|---|
| P50 in-VPC LB latency | 2.63 msmeasured (CWM Bench) | 2.50 ms | -4.9% | 95% |
| P95 in-VPC LB latency | 4.53 msmeasured (CWM Bench) | 4.40 ms | -2.9% | 97% |
| App-server CPU | 0.48%measured (CWM Bench) | 0.39% | -18.3% | 82% |
| Throughput | 8.87 req/smeasured (CWM Bench) | 8.87 req/s | +0.0% | 100% |
| Error RateAbsolute difference in percentage points; see tolerance rule above. | 0.00%measured (CWM Bench) | 0.00% | 0.0000000 pp | 100% |
| Cost / Hour | $0.4545 $/hrpublic price list (not a measured bill) | $0.4556 $/hr | +0.2% | 100% |
| Owned metric | Value | Meaning |
|---|---|---|
| Target RPS | 10 | target load |
| Goodput | 8.874 RPS | warmup + steady window |
| P50 latency | 2.628 ms | k6 in-VPC, internal LB; no internet travel |
| P95 latency | 4.530 ms | k6 in-VPC, internal LB; no internet travel |
| P99 latency | 7.238 ms | k6 in-VPC, internal LB; no internet travel |
| App CPU | 0.48% | owned application host metric |
| DB CPU | 3.34% | owned database metric |
| Derived CPU blend | 1.43% | (2 × app + DB) / 3; not measured app CPU |
| DB connections | 2 / ~500 | maximum observed / documented ceiling |
| CRUD errors | 0.00% | owned canonical CRUD mix |
| Cost / hour | $0.4545 | public price list |
Source: company-owned cwm-bench campaign. CPU score uses app hosts only; database CPU and the derived blend are diagnostics. Latency score uses the measured in-VPC internal-LB path.
| P99 diagnostic (not scored) | Measured (cwm-bench) | Engine prediction | Provenance |
|---|---|---|---|
| P99 latency | 7.238 ms | 7.025 ms | Owned fit run · prediction shown for diagnostic comparison only; excluded from scoring. |
| Metric | Reference scored source | Simulated | Delta | Accuracy |
|---|---|---|---|---|
| P50 in-VPC LB latency | 2.19 msmeasured (CWM Bench) | 2.40 ms | +9.7% | 90% |
| P95 in-VPC LB latency | 4.03 msmeasured (CWM Bench) | 4.25 ms | +5.5% | 95% |
| App-server CPU | 1.79%measured (CWM Bench) | 1.90% | +6.0% | 94% |
| Throughput | 87.62 req/smeasured (CWM Bench) | 87.62 req/s | +0.0% | 100% |
| Error RateAbsolute difference in percentage points; see tolerance rule above. | 0.00%measured (CWM Bench) | 0.00% | 0.0000000 pp | 100% |
| Cost / Hour | $0.4545 $/hrpublic price list (not a measured bill) | $0.4580 $/hr | +0.8% | 99% |
| Owned metric | Value | Meaning |
|---|---|---|
| Target RPS | 100 | target load |
| Goodput | 87.625 RPS | warmup + steady window |
| P50 latency | 2.188 ms | k6 in-VPC, internal LB; no internet travel |
| P95 latency | 4.028 ms | k6 in-VPC, internal LB; no internet travel |
| P99 latency | 6.787 ms | k6 in-VPC, internal LB; no internet travel |
| App CPU | 1.79% | owned application host metric |
| DB CPU | 4.43% | owned database metric |
| Derived CPU blend | 2.67% | (2 × app + DB) / 3; not measured app CPU |
| DB connections | 8 / ~500 | maximum observed / documented ceiling |
| CRUD errors | 0.00% | owned canonical CRUD mix |
| Cost / hour | $0.4545 | public price list |
Source: company-owned cwm-bench campaign. CPU score uses app hosts only; database CPU and the derived blend are diagnostics. Latency score uses the measured in-VPC internal-LB path.
| P99 diagnostic (not scored) | Measured (cwm-bench) | Engine prediction | Provenance |
|---|---|---|---|
| P99 latency | 6.787 ms | 7.049 ms | Owned fit run · prediction shown for diagnostic comparison only; excluded from scoring. |
| Metric | Reference scored source | Simulated | Delta | Accuracy |
|---|---|---|---|---|
| P50 in-VPC LB latency | 2.00 msmeasured (CWM Bench) | 2.00 ms | +0.2% | 100% |
| P95 in-VPC LB latency | 3.76 msmeasured (CWM Bench) | 3.75 ms | -0.3% | 100% |
| App-server CPU | 8.63%measured (CWM Bench) | 8.61% | -0.2% | 100% |
| Throughput | 437.62 req/smeasured (CWM Bench) | 437.62 req/s | +0.0% | 100% |
| Error RateAbsolute difference in percentage points; see tolerance rule above. | 0.00%measured (CWM Bench) | 0.00% | 0.0000000 pp | 100% |
| Cost / Hour | $0.4545 $/hrpublic price list (not a measured bill) | $0.4511 $/hr | -0.8% | 99% |
| Owned metric | Value | Meaning |
|---|---|---|
| Target RPS | 500 | target load |
| Goodput | 437.623 RPS | warmup + steady window |
| P50 latency | 1.995 ms | k6 in-VPC, internal LB; no internet travel |
| P95 latency | 3.760 ms | k6 in-VPC, internal LB; no internet travel |
| P99 latency | 7.205 ms | k6 in-VPC, internal LB; no internet travel |
| App CPU | 8.63% | owned application host metric |
| DB CPU | 8.22% | owned database metric |
| Derived CPU blend | 8.49% | (2 × app + DB) / 3; not measured app CPU |
| DB connections | 43 / ~500 | maximum observed / documented ceiling |
| CRUD errors | 0.00% | owned canonical CRUD mix |
| Cost / hour | $0.4545 | public price list |
Source: company-owned cwm-bench campaign. CPU score uses app hosts only; database CPU and the derived blend are diagnostics. Latency score uses the measured in-VPC internal-LB path.
| P99 diagnostic (not scored) | Measured (cwm-bench) | Engine prediction | Provenance |
|---|---|---|---|
| P99 latency | 7.205 ms | 7.157 ms | Owned fit run · prediction shown for diagnostic comparison only; excluded from scoring. |
Holdout runs (not fit inputs)
b2f414f2d8ec0e72c4ee980baff0a0bff3a4903c (seed 20240601); current predictions come from this live benchmark response (seed 20240601). None of these holdouts was used to fit the CPU or latency coefficients.Canonical AWS 1,000-RPS Burst; seeded engine mean app-host CPU and modeled in-VPC internal-LB latency before owned CPU/latency calibration. Cost, when scored, is an AWS us-east-2 price-list reference, not a measurement. Documentation latency and CPU figures are context only, not score inputs.
| Holdout / metric | Observed (cwm-bench) | Historical prediction (pinned revision) | Current prediction (live) |
|---|---|---|---|
| Burst · us-east-2 · 1,000 RPS App-server CPU | 13.36% | 95.10% | 17.00% |
| Burst · us-east-2 · 1,000 RPS P50 in-VPC internal-LB latency | 1.93 ms | 70.00 ms | 1.45 ms |
| Burst · us-east-2 · 1,000 RPS P95 in-VPC internal-LB latency | 3.71 ms | 159.00 ms | 3.10 ms |
| Burst · us-east-2 · 1,000 RPS Whole-run goodput | 875.12 RPS | 875.12 RPS | 875.12 RPS |
| Burst · us-east-2 · 1,000 RPS CRUD error rate | 0.0000952% | 0.00% | 0.00% |
| Later-day · us-east-2 · 100 RPS App-server CPU | ~1.68% | No separate prediction for this run | No separate prediction for this run |
| Later-day · us-east-2 · 100 RPS P50 in-VPC internal-LB latency | ~2.31 ms | No separate prediction for this run | No separate prediction for this run |
| Later-day · us-east-2 · 100 RPS P95 in-VPC internal-LB latency | ~4.27 ms | No separate prediction for this run | No separate prediction for this run |
| Later-day · us-east-2 · 100 RPS CRUD error rate | ~0.00% | No separate prediction for this run | No separate prediction for this run |
| Second region · us-west-2 · 100 RPS App-server CPU | ~1.57% | No separate prediction for this run | No separate prediction for this run |
| Second region · us-west-2 · 100 RPS P50 in-VPC internal-LB latency | ~2.00 ms | No separate prediction for this run | No separate prediction for this run |
| Second region · us-west-2 · 100 RPS P95 in-VPC internal-LB latency | ~3.84 ms | No separate prediction for this run | No separate prediction for this run |
| Second region · us-west-2 · 100 RPS CRUD error rate | ~0.00% | No separate prediction for this run | No separate prediction for this run |
On the scored Burst holdout, current app CPU is 17.00% versus measured 13.36%; P50 is 1.45 ms versus 1.93 ms, and P95 is 3.10 ms versus 3.71 ms. Remaining gaps above 10% relative error: P50 in-VPC LB latency (24.9%), P95 in-VPC LB latency (16.5%), App-server CPU (27.2%). The later-day and second-region observations have no independent engine runs and do not contribute to the composite.
| Metric | Reference scored source | Simulated | Delta | Accuracy |
|---|---|---|---|---|
| P50 in-VPC LB latency | 1.93 msmeasured (CWM Bench) | 1.45 ms | -24.9% | 75% |
| P95 in-VPC LB latency | 3.71 msmeasured (CWM Bench) | 3.10 ms | -16.5% | 84% |
| App-server CPU | 13.36%measured (CWM Bench) | 17.00% | +27.2% | 73% |
| Throughput | 875.12 req/smeasured (CWM Bench) | 875.12 req/s | +0.0% | 100% |
| Error RateAbsolute difference in percentage points; see tolerance rule above. | 0.0000952%measured (CWM Bench) | 0.00% | 0.0000952 pp | 100% |
| Cost / Hour | $0.4545 $/hrpublic price list (not a measured bill) | $0.4523 $/hr | -0.5% | 100% |
| Owned metric | Value | Meaning |
|---|---|---|
| Target RPS | 1,000 | target load |
| Goodput | 875.116 RPS | warmup + steady window |
| P50 latency | 1.931 ms | k6 in-VPC, internal LB; no internet travel |
| P95 latency | 3.714 ms | k6 in-VPC, internal LB; no internet travel |
| P99 latency | 7.017 ms | k6 in-VPC, internal LB; no internet travel |
| App CPU | 13.36% | owned application host metric |
| DB CPU | 10.91% | owned database metric |
| Derived CPU blend | 12.54% | (2 × app + DB) / 3; not measured app CPU |
| DB connections | 55 / ~500 | maximum observed / documented ceiling |
| CRUD errors | 0.0000952% | owned canonical CRUD mix |
| Cost / hour | $0.4545 | public price list |
Source: company-owned cwm-bench campaign. CPU score uses app hosts only; database CPU and the derived blend are diagnostics. Latency score uses the measured in-VPC internal-LB path. The measured error rate is 0.0000952% and whole-run goodput is 875.12 RPS, including ramp-up. No dropped requests are implied.
| P99 diagnostic (not scored) | Measured (cwm-bench) | Engine prediction | Provenance |
|---|---|---|---|
| P99 latency | 7.017 ms | 7.292 ms | Owned holdout run · prediction shown for diagnostic comparison only; excluded from scoring. |
Later-day · us-east-2 · 100 RPS
~1.68% app CPU · ~2.31 ms median · ~4.27 ms P95 · 0% errors
No separate prediction for this run; canonical 100-RPS predictions are not reused.
Second region · us-west-2 · 100 RPS
~1.57% app CPU · ~2.00 ms median · ~3.84 ms P95 · 0% errors
No separate prediction for this run; canonical 100-RPS predictions are not reused.
Pricing accuracy is validated continuously via automated drift checks in CI. See the Simulation Fidelity page for per-provider cost benchmark details across all five cloud providers.
Want to benchmark a different architecture? Open the Workspace to build and simulate any topology.
Try the Simulation Yourself
Load the same 3-tier AWS scenario in the interactive workspace and compare what you see with the reference values on this page.
