Cloud World Model on Smithery

    Simulation Accuracy Benchmark

    Lean · Overall: 93.8% accurate
    Lean burst holdout (not used for tuning): 85.4% · 3 tuning runs + 1 holdout
    Typical 300 RPS holdout: 24.3/100 before fit; 84.2/100 after fit
    Typical 500 RPS diagnostic, extrapolated beyond fitted range, not in the headline: 71.0/100
    Measured (CWM Bench)

    This page publishes a reproducible, company-owned cwm-bench measurement of the canonical AWS architecture and clearly separates it from engine predictions and provider-documentation references.

    See also: Simulation Fidelity — benchmark data and accuracy ranges for all five cloud providers.

    AWS scored app CPU, in-VPC internal-load-balancer latency, goodput and CRUD errors come from the pinned owned cwm-bench campaign. Cost uses the us-east-2 price list. Tuning uses 10 / 100 / 500 RPS; 1,000 RPS, later-day and us-west-2 are holdouts. See the Methodology section below for full source citations.

    GCP, Azure, OCI, and DigitalOcean scores use documentation references, not owned measurements. They are shown separately from canonical AWS and are not a cross-provider ranking.

    OpenShift reference scenarios
    Reference-only
    OpenShift is a platform overlay, not a sixth cloud provider, so it is not included in the measured provider scorecard yet.
    Cost: Estimated
    Latency, CPU, throughput, errors: Extrapolated

    Independently sourced coverage currently includes rosa-hcp, rosa-classic, aro, openshift-dedicated, self-managed across 6 region-specific scenarios. Platform fees are estimated from published product/pricing information; performance behavior is not a topology-matched public load test.

    View source-backed OpenShift reference output
    All Providers at a Glance
    Scores are grouped by evidence basis, not ranked together. Documentation references have not been checked against owned measurements. Click a row to jump to its detailed results below.
    ProviderScoring basisOverall
    Owned lean-weight measurement comparison (cost: public price list); typical holdout shown separately below
    AWS (lean)
    Measured (CWM Bench)Measured performance and goodput/error; cost is a public price-list reference, not a measured bill.93.8%
    Documentation-reference comparisons — not yet measured
    GCP
    Documentation reference (not yet measured)Documentation references have not been checked against owned measurements.98.0%
    Azure
    Documentation reference (not yet measured)Documentation references have not been checked against owned measurements.97.8%
    Oracle Cloud (OCI)
    Documentation reference (not yet measured)Documentation references have not been checked against owned measurements.96.2%
    DigitalOcean
    Documentation reference (not yet measured)Documentation references have not been checked against owned measurements.97.4%

    Click any row to load that provider's full benchmark results below.

    CPU Accuracy by Scenario
    Simulated vs. reference CPU utilization at each traffic scenario. CPU is one of the largest single drivers of each provider's score, so this shows why a provider lands where it does — and makes future calibration changes visible at a glance. Each cell colors the simulated value by how closely it tracks the reference (green ≤ 10%, yellow ≤ 25%, red above 25% off).
    Provider / scoring basis
    Idle
    10 req/s
    Normal
    100 req/s
    Peak
    500 req/s
    Burst
    1,000 req/s
    Owned measurement comparison
    AWSMeasured (CWM Bench)
    0.4%measured app CPU 0.5%
    1.9%measured app CPU 1.8%
    8.6%measured app CPU 8.6%
    17.0%measured app CPU 13.4%
    Documentation references — not checked against owned measurements
    GCPDocumentation reference (not yet measured)
    6.3%doc ref 6.0%
    22.0%doc ref 22.0%
    57.3%doc ref 58.0%
    88.1%doc ref 84.0%
    AzureDocumentation reference (not yet measured)
    5.4%doc ref 5.0%
    18.5%doc ref 19.0%
    52.3%doc ref 54.0%
    86.1%doc ref 81.0%
    Oracle Cloud (OCI)Documentation reference (not yet measured)
    4.3%doc ref 4.0%
    12.1%doc ref 12.0%
    44.4%doc ref 48.0%
    82.6%doc ref 78.0%
    DigitalOceanDocumentation reference (not yet measured)
    6.3%doc ref 6.0%
    15.4%doc ref 15.0%
    41.9%doc ref 44.0%
    70.3%doc ref 68.0%

    AWS cells show predicted application-server CPU against measured application-server CPU; other providers show simulated output against documentation-sourced references. Click any row to load that provider's full benchmark below.

    6th-Gen AWS Instance Accuracy
    Per-instance benchmark scores for six 6th-generation AWS instance types (m6i, c6i, r6i) across all four traffic scenarios. Scores are derived from the same reference scenarios used by the canonical AWS benchmark above.
    InstanceOverall
    m6i.large96.8%
    m6i.xlarge99.2%
    c6i.large97.8%
    c6i.xlarge99.2%
    r6i.large96.1%
    r6i.xlarge98.1%

    Each row benchmarks the simulator against doc-sourced reference values for that specific instance type. Cost, Latency, and Perf are category averages across the four traffic scenarios.

    Select a provider to load its canonical benchmark automatically. Measured (CWM Bench) — Measured performance and goodput/error; cost is a public price-list reference, not a measured bill.

    AWS canonical comparison scores owned cwm-bench app-host CPU and k6 in-VPC internal-LB latency against seeded engine predictions, with owned goodput/errors and AWS price-list cost. Idle/Normal/Peak are fit rungs; Burst is a holdout. Later-day and second-region have no separate predictions.

    Lean Overall Simulation Accuracy: 93.8%

    Weighted composite across P50/P95 latency, CPU utilization, throughput, error rate, and cost — averaged over four traffic scenarios (Idle, Normal, Peak, Burst).

    Lean burst holdout (not used for tuning): 85.4% · 3 tuning runs + 1 holdout

    This 93.8% overall result and burst holdout measure lean app weight. Self-serve and the GitHub App default to typical; its separate owned 300 RPS holdout below is not included in the lean overall score.

    Throughput largely follows the prescribed workload; cost uses the AWS public price list, not a measured bill.

    Measured performance and goodput/error; cost is a public price-list reference, not a measured bill.

    Idle

    95%

    Normal

    95%

    Peak

    100%

    Burst

    85%

    Typical app weight (owned benchmark)
    Separate frozen before- and after-fit comparison for the exact owned typical setup; not part of the lean overall score.

    300 RPS holdout — before fit: 24.3/100 (27.0 without cost)

    300 RPS independent holdout — after fit: 84.2/100 (82.4 without cost)

    Engine 1.2.6; aws-crud/typical:typical-v1-20260927c:6aa574d7ff9d3080b88b221bcd59f7d218ae37f0. Only typical-fit-20, typical-fit-100, typical-fit-200 were used for fitting. Neither 300 nor 500 RPS was used to tune the fit.

    MetricPredictedMeasured / referenceMetric scoreWeightWeighted contribution
    P504.17 ms4.62 ms90.10.2018.0
    P9512.78 ms26.92 ms47.50.2511.9
    App CPU14.74 %15.30 %96.30.2019.3
    Throughput262.60 RPS262.61 RPS100.00.1515.0
    Error rate0.0000 %0.0044 %100.00.1010.0
    Cost0.454581 USD/hour0.4545 USD/hour (reference)100.00.1010.0

    Throughput: whole-run goodput (5 min warmup plus 15 min steady); the prediction is a point rate.

    Cost: Cost was not measured. USD 0.4545/hour is a public list-price reference, not a bill. The post-fit scored prediction excludes modeled egress; the historical before-fit prediction included it, so their cost scores are not on identical bases.

    500 RPS — diagnostic, extrapolated beyond fitted range, not in the headline: 71.0/100 (67.9 without cost)

    Engine 1.2.6; aws-crud/typical:typical-v1-20260927c:6aa574d7ff9d3080b88b221bcd59f7d218ae37f0. Only typical-fit-20, typical-fit-100, typical-fit-200 were used for fitting. This point was not used to tune the fit; the 300 RPS result remains an independent holdout.

    MetricPredictedMeasured / referenceMetric score
    P503.451 ms5.188 ms66.5
    P9516.333 ms88.190 ms18.5
    App CPU24.264 %26.342 %92.1
    Throughput437.588 RPS437.577 RPS100.0
    Error rate0.000 %0.010 %97.2
    Cost0.451107 USD/hour0.4545 USD/hour (reference)99.3

    Throughput: whole-run goodput (5 min warmup plus 15 min steady); the prediction is a point rate. Cost: USD 0.4545/hour is a no-egress public list-price reference, not a measured bill; scored prediction excludes modeled egress.

    One campaign, one day, us-east-2, one attempt per rung; engine 1.2.3 predictions frozen before measurement (seed 20240601, step 6). Measured cost missing. Internal ALB + 2 × m5.large + db.r5.large MySQL 8.0; two workers and pool 250 per server.

    MetricPredictedMeasured / referenceMetric scoreWeightWeighted contribution
    P5017 ms4.62 ms0.00.200.0
    P9542 ms26.92 ms44.00.2511.0
    App CPU59.553 %15.30 %0.00.200.0
    Throughput299 RPS262.61 RPS86.10.1512.9
    Error rate0.284 %0.0044 %3.60.100.4
    Cost0.94 USD/hour0.4545 USD/hour (reference)0.00.100.0

    P99 (unweighted context, not part of the score): 236 vs 97.5 ms; score 0.

    Throughput: whole-run goodput (5 min warmup plus 15 min steady); the prediction is a point rate.

    Cost: Cost was not measured. Score uses the list-price reference of USD 0.4545 per hour (ALB 0.0225 + 2 × 0.096 + 0.24, no egress, no generator). Reference, not a bill.

    Optional, unscored steady-window context: 300.0 RPS gives throughput 99.7 and total 26.3.

    Historical pre-fit context — fit data, not independent holdout accuracy; old saturation diagnostic is not the extrapolated score above

    • fit-20 (fit data, not holdout): 23.0 (15.3 without cost)
    • fit-100 (fit data, not holdout): 19.8 (14.9 without cost)
    • fit-200 (fit data, not holdout): 16.1 (14.8 without cost)
    • saturation-500 (diagnostic only): 25.6 (28.5 without cost)

    Sources: cwm-bench typical report · summary export · scores · frozen predictions · measurement commit 6aa574d7ff9d

    Tune the Architecture
    Choose a cloud provider then adjust instance counts and types to see how simulation accuracy changes for your setup. Reference values stay fixed at the canonical AWS architecture — you're exploring how the simulator responds across providers and resource sizes.
    2
    110

    Results update the architecture diagram and scenario tables below.

    Canonical Architecture

    AWS 3-Tier Web App
    us-east-2
    Standard three-tier web application: Application Load Balancer distributing traffic across two m5.large EC2 instances backed by a single db.r5.large RDS MySQL instance (Single-AZ, us-east-2).
    RoleInstance TypeCountCost / hr
    Load Balancer
    ALB (Application Load Balancer)1$0.0225
    Web / App Server
    m5.large2$0.1920
    Database
    db.r5.large (MySQL 8.0, Single-AZ)1$0.2400
    Total (official AWS pricing, us-east-2)$0.4545/hr

    Want to run this exact scenario yourself? Open the Simulation Workspace and configure the same resources.

    Traffic Scenario Comparisons

    Each table shows Reference vs. Simulated values, the percentage delta, and an accuracy badge (green ≥ 90%, yellow ≥ 75%, red below 75%).

    Error-rate accuracy: absolute simulated-versus-reference difference ≤ 0.01 percentage points scores 100%. Above that, the score is the higher of 100 × 0.01 ÷ difference (in percentage points) and the usual relative-accuracy score (100 minus absolute percent difference, floored at 0; zero if the reference is zero and the prediction is not).

    Tuning runs (fit)

    Idle
    10 req/s
    Composite score:94.6%
    Minimal background traffic — keep-alive checks, health probes, and occasional real requests.
    MetricReference scored sourceSimulatedDeltaAccuracy
    P50 in-VPC LB latency2.63 msmeasured (CWM Bench)2.50 ms-4.9%
    95%
    P95 in-VPC LB latency4.53 msmeasured (CWM Bench)4.40 ms-2.9%
    97%
    App-server CPU0.48%measured (CWM Bench)0.39%-18.3%
    82%
    Throughput8.87 req/smeasured (CWM Bench)8.87 req/s+0.0%
    100%
    Error RateAbsolute difference in percentage points; see tolerance rule above.0.00%measured (CWM Bench)0.00%0.0000000 pp
    100%
    Cost / Hour$0.4545 $/hrpublic price list (not a measured bill)$0.4556 $/hr+0.2%
    100%
    Idle — Owned cwm-bench measurements
    10 RPS target
    fit measurement
    Engine comparison available
    Minimal background traffic — keep-alive checks, health probes, and occasional real requests.
    Owned metricValueMeaning
    Target RPS10target load
    Goodput8.874 RPSwarmup + steady window
    P50 latency2.628 msk6 in-VPC, internal LB; no internet travel
    P95 latency4.530 msk6 in-VPC, internal LB; no internet travel
    P99 latency7.238 msk6 in-VPC, internal LB; no internet travel
    App CPU0.48%owned application host metric
    DB CPU3.34%owned database metric
    Derived CPU blend1.43%(2 × app + DB) / 3; not measured app CPU
    DB connections2 / ~500maximum observed / documented ceiling
    CRUD errors0.00%owned canonical CRUD mix
    Cost / hour$0.4545public price list

    Source: company-owned cwm-bench campaign. CPU score uses app hosts only; database CPU and the derived blend are diagnostics. Latency score uses the measured in-VPC internal-LB path.

    P99 diagnostic (not scored)Measured (cwm-bench)Engine predictionProvenance
    P99 latency7.238 ms7.025 msOwned fit run · prediction shown for diagnostic comparison only; excluded from scoring.
    Normal
    100 req/s
    Composite score:95.4%
    Typical business-hours traffic with steady sustained load.
    MetricReference scored sourceSimulatedDeltaAccuracy
    P50 in-VPC LB latency2.19 msmeasured (CWM Bench)2.40 ms+9.7%
    90%
    P95 in-VPC LB latency4.03 msmeasured (CWM Bench)4.25 ms+5.5%
    95%
    App-server CPU1.79%measured (CWM Bench)1.90%+6.0%
    94%
    Throughput87.62 req/smeasured (CWM Bench)87.62 req/s+0.0%
    100%
    Error RateAbsolute difference in percentage points; see tolerance rule above.0.00%measured (CWM Bench)0.00%0.0000000 pp
    100%
    Cost / Hour$0.4545 $/hrpublic price list (not a measured bill)$0.4580 $/hr+0.8%
    99%
    Normal — Owned cwm-bench measurements
    100 RPS target
    fit measurement
    Engine comparison available
    Typical business-hours traffic with steady sustained load.
    Owned metricValueMeaning
    Target RPS100target load
    Goodput87.625 RPSwarmup + steady window
    P50 latency2.188 msk6 in-VPC, internal LB; no internet travel
    P95 latency4.028 msk6 in-VPC, internal LB; no internet travel
    P99 latency6.787 msk6 in-VPC, internal LB; no internet travel
    App CPU1.79%owned application host metric
    DB CPU4.43%owned database metric
    Derived CPU blend2.67%(2 × app + DB) / 3; not measured app CPU
    DB connections8 / ~500maximum observed / documented ceiling
    CRUD errors0.00%owned canonical CRUD mix
    Cost / hour$0.4545public price list

    Source: company-owned cwm-bench campaign. CPU score uses app hosts only; database CPU and the derived blend are diagnostics. Latency score uses the measured in-VPC internal-LB path.

    P99 diagnostic (not scored)Measured (cwm-bench)Engine predictionProvenance
    P99 latency6.787 ms7.049 msOwned fit run · prediction shown for diagnostic comparison only; excluded from scoring.
    Peak
    500 req/s
    Composite score:99.8%
    Peak business load — product launch, end-of-day batch, or marketing campaign spike.
    MetricReference scored sourceSimulatedDeltaAccuracy
    P50 in-VPC LB latency2.00 msmeasured (CWM Bench)2.00 ms+0.2%
    100%
    P95 in-VPC LB latency3.76 msmeasured (CWM Bench)3.75 ms-0.3%
    100%
    App-server CPU8.63%measured (CWM Bench)8.61%-0.2%
    100%
    Throughput437.62 req/smeasured (CWM Bench)437.62 req/s+0.0%
    100%
    Error RateAbsolute difference in percentage points; see tolerance rule above.0.00%measured (CWM Bench)0.00%0.0000000 pp
    100%
    Cost / Hour$0.4545 $/hrpublic price list (not a measured bill)$0.4511 $/hr-0.8%
    99%
    Peak — Owned cwm-bench measurements
    500 RPS target
    fit measurement
    Engine comparison available
    Peak business load — product launch, end-of-day batch, or marketing campaign spike.
    Owned metricValueMeaning
    Target RPS500target load
    Goodput437.623 RPSwarmup + steady window
    P50 latency1.995 msk6 in-VPC, internal LB; no internet travel
    P95 latency3.760 msk6 in-VPC, internal LB; no internet travel
    P99 latency7.205 msk6 in-VPC, internal LB; no internet travel
    App CPU8.63%owned application host metric
    DB CPU8.22%owned database metric
    Derived CPU blend8.49%(2 × app + DB) / 3; not measured app CPU
    DB connections43 / ~500maximum observed / documented ceiling
    CRUD errors0.00%owned canonical CRUD mix
    Cost / hour$0.4545public price list

    Source: company-owned cwm-bench campaign. CPU score uses app hosts only; database CPU and the derived blend are diagnostics. Latency score uses the measured in-VPC internal-LB path.

    P99 diagnostic (not scored)Measured (cwm-bench)Engine predictionProvenance
    P99 latency7.205 ms7.157 msOwned fit run · prediction shown for diagnostic comparison only; excluded from scoring.

    Holdout runs (not fit inputs)

    Holdouts — before and after calibration
    Observations are from the pinned cwm-bench campaign. Historical predictions are a snapshot from pre-calibration engine revision b2f414f2d8ec0e72c4ee980baff0a0bff3a4903c (seed 20240601); current predictions come from this live benchmark response (seed 20240601). None of these holdouts was used to fit the CPU or latency coefficients.

    Canonical AWS 1,000-RPS Burst; seeded engine mean app-host CPU and modeled in-VPC internal-LB latency before owned CPU/latency calibration. Cost, when scored, is an AWS us-east-2 price-list reference, not a measurement. Documentation latency and CPU figures are context only, not score inputs.

    Holdout / metricObserved (cwm-bench)Historical prediction (pinned revision)Current prediction (live)
    Burst · us-east-2 · 1,000 RPS
    App-server CPU
    13.36%95.10%17.00%
    Burst · us-east-2 · 1,000 RPS
    P50 in-VPC internal-LB latency
    1.93 ms70.00 ms1.45 ms
    Burst · us-east-2 · 1,000 RPS
    P95 in-VPC internal-LB latency
    3.71 ms159.00 ms3.10 ms
    Burst · us-east-2 · 1,000 RPS
    Whole-run goodput
    875.12 RPS875.12 RPS875.12 RPS
    Burst · us-east-2 · 1,000 RPS
    CRUD error rate
    0.0000952%0.00%0.00%
    Later-day · us-east-2 · 100 RPS
    App-server CPU
    ~1.68%No separate prediction for this runNo separate prediction for this run
    Later-day · us-east-2 · 100 RPS
    P50 in-VPC internal-LB latency
    ~2.31 msNo separate prediction for this runNo separate prediction for this run
    Later-day · us-east-2 · 100 RPS
    P95 in-VPC internal-LB latency
    ~4.27 msNo separate prediction for this runNo separate prediction for this run
    Later-day · us-east-2 · 100 RPS
    CRUD error rate
    ~0.00%No separate prediction for this runNo separate prediction for this run
    Second region · us-west-2 · 100 RPS
    App-server CPU
    ~1.57%No separate prediction for this runNo separate prediction for this run
    Second region · us-west-2 · 100 RPS
    P50 in-VPC internal-LB latency
    ~2.00 msNo separate prediction for this runNo separate prediction for this run
    Second region · us-west-2 · 100 RPS
    P95 in-VPC internal-LB latency
    ~3.84 msNo separate prediction for this runNo separate prediction for this run
    Second region · us-west-2 · 100 RPS
    CRUD error rate
    ~0.00%No separate prediction for this runNo separate prediction for this run

    On the scored Burst holdout, current app CPU is 17.00% versus measured 13.36%; P50 is 1.45 ms versus 1.93 ms, and P95 is 3.10 ms versus 3.71 ms. Remaining gaps above 10% relative error: P50 in-VPC LB latency (24.9%), P95 in-VPC LB latency (16.5%), App-server CPU (27.2%). The later-day and second-region observations have no independent engine runs and do not contribute to the composite.

    Burst
    1,000 req/s
    Composite score:85.4%
    Traffic burst exceeding normal peak — flash sale, viral event, or coordinated load test.
    MetricReference scored sourceSimulatedDeltaAccuracy
    P50 in-VPC LB latency1.93 msmeasured (CWM Bench)1.45 ms-24.9%
    75%
    P95 in-VPC LB latency3.71 msmeasured (CWM Bench)3.10 ms-16.5%
    84%
    App-server CPU13.36%measured (CWM Bench)17.00%+27.2%
    73%
    Throughput875.12 req/smeasured (CWM Bench)875.12 req/s+0.0%
    100%
    Error RateAbsolute difference in percentage points; see tolerance rule above.0.0000952%measured (CWM Bench)0.00%0.0000952 pp
    100%
    Cost / Hour$0.4545 $/hrpublic price list (not a measured bill)$0.4523 $/hr-0.5%
    100%
    Burst — Owned cwm-bench measurements
    1,000 RPS target
    holdout measurement
    Engine comparison available
    Traffic burst exceeding normal peak — flash sale, viral event, or coordinated load test.
    Owned metricValueMeaning
    Target RPS1,000target load
    Goodput875.116 RPSwarmup + steady window
    P50 latency1.931 msk6 in-VPC, internal LB; no internet travel
    P95 latency3.714 msk6 in-VPC, internal LB; no internet travel
    P99 latency7.017 msk6 in-VPC, internal LB; no internet travel
    App CPU13.36%owned application host metric
    DB CPU10.91%owned database metric
    Derived CPU blend12.54%(2 × app + DB) / 3; not measured app CPU
    DB connections55 / ~500maximum observed / documented ceiling
    CRUD errors0.0000952%owned canonical CRUD mix
    Cost / hour$0.4545public price list

    Source: company-owned cwm-bench campaign. CPU score uses app hosts only; database CPU and the derived blend are diagnostics. Latency score uses the measured in-VPC internal-LB path. The measured error rate is 0.0000952% and whole-run goodput is 875.12 RPS, including ramp-up. No dropped requests are implied.

    P99 diagnostic (not scored)Measured (cwm-bench)Engine predictionProvenance
    P99 latency7.017 ms7.292 msOwned holdout run · prediction shown for diagnostic comparison only; excluded from scoring.
    Additional holdouts — observations only
    Approximate readings from the pinned campaign export. These runs were not fit inputs or scored; no separate engine prediction was made for either run.

    Later-day · us-east-2 · 100 RPS

    ~1.68% app CPU · ~2.31 ms median · ~4.27 ms P95 · 0% errors

    No separate prediction for this run; canonical 100-RPS predictions are not reused.

    Second region · us-west-2 · 100 RPS

    ~1.57% app CPU · ~2.00 ms median · ~3.84 ms P95 · 0% errors

    No separate prediction for this run; canonical 100-RPS predictions are not reused.

    Methodology & Reference Sources
    How reference values are derived and what each data source covers.

    Pricing accuracy is validated continuously via automated drift checks in CI. See the Simulation Fidelity page for per-provider cost benchmark details across all five cloud providers.

    Want to benchmark a different architecture? Open the Workspace to build and simulate any topology.

    Try the Simulation Yourself

    Load the same 3-tier AWS scenario in the interactive workspace and compare what you see with the reference values on this page.