Cloud World Model on Smithery

    Simulation Accuracy Benchmark

    Overall: 97.4% accurate

    This page publishes a reproducible, company-owned cwm-bench measurement of the canonical AWS architecture and clearly separates it from engine predictions and provider-documentation references.

    See also: Simulation Fidelity — benchmark data and accuracy ranges for all five cloud providers.

    AWS goodput, CRUD errors, app-local latency, role-specific CPU, and connection diagnostics come from the pinned owned cwm-bench campaign. Scored latency and blended-CPU references remain documentation-backed. The fit split uses 10 / 100 / 500 RPS; 1,000 RPS burst is a holdout. See the Methodology section below for full source citations.

    OpenShift reference scenarios
    Reference-only
    OpenShift is a platform overlay, not a sixth cloud provider, so it is not included in the measured provider scorecard yet.
    Cost: Estimated
    Latency, CPU, throughput, errors: Extrapolated

    Independently sourced coverage currently includes rosa-hcp, rosa-classic, aro, openshift-dedicated, self-managed across 6 region-specific scenarios. Platform fees are estimated from published product/pricing information; performance behavior is not a topology-matched public load test.

    View source-backed OpenShift reference output
    All Providers at a Glance
    Overall accuracy score and per-scenario scores for all five providers. Click a row to jump to its detailed results below.
    ProviderOverall
    AWS
    97.4%
    GCP
    97.9%
    Azure
    97.8%
    Oracle Cloud (OCI)
    96.2%
    DigitalOcean
    97.4%

    Click any row to load that provider's full benchmark results below.

    CPU Accuracy by Scenario
    Simulated vs. reference CPU utilization at each traffic scenario. CPU is one of the largest single drivers of each provider's score, so this shows why a provider lands where it does — and makes future calibration changes visible at a glance. Each cell colors the simulated value by how closely it tracks the reference (green ≤ 10%, yellow ≤ 25%, red above 25% off).
    Provider
    Idle
    10 req/s
    Normal
    100 req/s
    Peak
    500 req/s
    Burst
    1,000 req/s
    AWS
    4.9%doc blended ref 5%owned app CPU 0.5%
    19.1%doc blended ref 20%owned app CPU 1.8%
    49.2%doc blended ref 52%owned app CPU 8.6%
    74.9%doc blended ref 78%owned app CPU 13.4%
    GCP
    6.3%doc ref 6%
    22.0%doc ref 22%
    57.3%doc ref 58%
    88.1%doc ref 84%
    Azure
    5.4%doc ref 5%
    18.5%doc ref 19%
    52.3%doc ref 54%
    86.1%doc ref 81%
    Oracle Cloud (OCI)
    4.3%doc ref 4%
    12.1%doc ref 12%
    44.4%doc ref 48%
    82.6%doc ref 78%
    DigitalOcean
    6.3%doc ref 6%
    15.4%doc ref 15%
    41.9%doc ref 44%
    70.3%doc ref 68%

    AWS cells show calibrated engine CPU against documentation-backed blended-CPU references, with the owned app CPU diagnostic shown separately; other providers show simulated output against documentation-sourced references. Click any row to load that provider's full benchmark below.

    6th-Gen AWS Instance Accuracy
    Per-instance benchmark scores for six 6th-generation AWS instance types (m6i, c6i, r6i) across all four traffic scenarios. Scores are derived from the same reference scenarios used by the canonical AWS benchmark above.
    InstanceOverall
    m6i.large96.8%
    m6i.xlarge99.2%
    c6i.large97.8%
    c6i.xlarge99.2%
    r6i.large96.1%
    r6i.xlarge98.1%

    Each row benchmarks the simulator against doc-sourced reference values for that specific instance type. Cost, Latency, and Perf are category averages across the four traffic scenarios.

    Select a provider to load its canonical benchmark automatically.

    AWS's canonical comparison uses calibrated engine output with owned cwm-bench goodput/error observations and documentation-backed end-to-end latency/blended-CPU references. App-local latency and role-split CPU are published diagnostics; Idle/normal/peak are fit rungs and Burst is a holdout.

    Overall Simulation Accuracy: 97.4%

    Weighted composite across P50/P95 latency, CPU utilization, throughput, error rate, and cost — averaged over four traffic scenarios (Idle, Normal, Peak, Burst).

    Idle

    95%

    Normal

    97%

    Peak

    99%

    Burst

    98%

    Tune the Architecture
    Choose a cloud provider then adjust instance counts and types to see how simulation accuracy changes for your setup. Reference values stay fixed at the canonical AWS architecture — you're exploring how the simulator responds across providers and resource sizes.
    2
    110

    Results update the architecture diagram and scenario tables below.

    Canonical Architecture

    AWS 3-Tier Web App
    us-east-2
    Standard three-tier web application: Application Load Balancer distributing traffic across two m5.large EC2 instances backed by a single db.r5.large RDS MySQL instance (Single-AZ, us-east-2).
    RoleInstance TypeCountCost / hr
    Load Balancer
    ALB (Application Load Balancer)1$0.0225
    Web / App Server
    m5.large2$0.1920
    Database
    db.r5.large (MySQL 8.0, Single-AZ)1$0.2400
    Total (official AWS pricing, us-east-2)$0.4545/hr

    Want to run this exact scenario yourself? Open the Simulation Workspace and configure the same resources.

    Traffic Scenario Comparisons

    Each table shows Reference vs. Simulated values, the percentage delta, and an accuracy badge (green ≥ 90%, yellow ≥ 75%, red below 75%).

    Idle
    10 req/s
    Composite score:95.1%
    Minimal background traffic — keep-alive checks, health probes, and occasional real requests.
    MetricReference scored sourceSimulatedDeltaAccuracy
    P50 Latency12 msdocumentation12 ms+0.0%
    100%
    P95 Latency28 msdocumentation33 ms+17.9%
    82%
    CPU Utilization5.00% %documentation4.90% %-2.0%
    98%
    Throughput8.87 req/sowned8.87 req/s+0.0%
    100%
    Error Rate0.00% %owned0.00% %+0.0%
    100%
    Cost / Hour$0.4545 $/hrprice list$0.4556 $/hr+0.2%
    100%
    Idle — Owned cwm-bench diagnostics
    10 RPS target
    fit measurement
    Engine comparison available
    Minimal background traffic — keep-alive checks, health probes, and occasional real requests.
    Owned metricValueMeaning
    Target RPS10target load
    Goodput8.874 RPSwarmup + steady window
    P50 latency2.628 msowned app-local measurement
    P95 latency4.530 msowned app-local measurement
    P99 latency7.238 msowned app-local measurement
    App CPU0.48%owned application host metric
    DB CPU3.34%owned database metric
    DB connections2 / ~500maximum observed / documented ceiling
    CRUD errors≈0%owned canonical CRUD mix
    Cost / hour$0.4545public price list

    Source: company-owned cwm-bench campaign. These are measurements, not simulated values. The scored comparison uses calibrated engine output with owned goodput/error and documentation-backed end-to-end latency/blended-CPU references.

    Normal
    100 req/s
    Composite score:97.4%
    Typical business-hours traffic with steady sustained load.
    MetricReference scored sourceSimulatedDeltaAccuracy
    P50 Latency18 msdocumentation17 ms-5.6%
    94%
    P95 Latency45 msdocumentation44 ms-2.2%
    98%
    CPU Utilization20.00% %documentation19.10% %-4.5%
    96%
    Throughput87.62 req/sowned87.62 req/s+0.0%
    100%
    Error Rate0.00% %owned0.00% %+0.0%
    100%
    Cost / Hour$0.4545 $/hrprice list$0.4580 $/hr+0.8%
    99%
    Normal — Owned cwm-bench diagnostics
    100 RPS target
    fit measurement
    Engine comparison available
    Typical business-hours traffic with steady sustained load.
    Owned metricValueMeaning
    Target RPS100target load
    Goodput87.625 RPSwarmup + steady window
    P50 latency2.188 msowned app-local measurement
    P95 latency4.028 msowned app-local measurement
    P99 latency6.787 msowned app-local measurement
    App CPU1.79%owned application host metric
    DB CPU4.43%owned database metric
    DB connections8 / ~500maximum observed / documented ceiling
    CRUD errors≈0%owned canonical CRUD mix
    Cost / hour$0.4545public price list

    Source: company-owned cwm-bench campaign. These are measurements, not simulated values. The scored comparison uses calibrated engine output with owned goodput/error and documentation-backed end-to-end latency/blended-CPU references.

    Peak
    500 req/s
    Composite score:98.8%
    Peak business load — product launch, end-of-day batch, or marketing campaign spike.
    MetricReference scored sourceSimulatedDeltaAccuracy
    P50 Latency38 msdocumentation38 ms+0.0%
    100%
    P95 Latency90 msdocumentation90 ms+0.0%
    100%
    CPU Utilization52.00% %documentation49.20% %-5.4%
    95%
    Throughput437.62 req/sowned437.62 req/s+0.0%
    100%
    Error Rate0.00% %owned0.00% %+0.0%
    100%
    Cost / Hour$0.4545 $/hrprice list$0.4511 $/hr-0.8%
    99%
    Peak — Owned cwm-bench diagnostics
    500 RPS target
    fit measurement
    Engine comparison available
    Peak business load — product launch, end-of-day batch, or marketing campaign spike.
    Owned metricValueMeaning
    Target RPS500target load
    Goodput437.623 RPSwarmup + steady window
    P50 latency1.995 msowned app-local measurement
    P95 latency3.760 msowned app-local measurement
    P99 latency7.205 msowned app-local measurement
    App CPU8.63%owned application host metric
    DB CPU8.22%owned database metric
    DB connections43 / ~500maximum observed / documented ceiling
    CRUD errors≈0%owned canonical CRUD mix
    Cost / hour$0.4545public price list

    Source: company-owned cwm-bench campaign. These are measurements, not simulated values. The scored comparison uses calibrated engine output with owned goodput/error and documentation-backed end-to-end latency/blended-CPU references.

    Burst
    1,000 req/s
    Composite score:98.4%
    Traffic burst exceeding normal peak — flash sale, viral event, or coordinated load test.
    MetricReference scored sourceSimulatedDeltaAccuracy
    P50 Latency68 msdocumentation70 ms+2.9%
    97%
    P95 Latency160 msdocumentation159 ms-0.6%
    99%
    CPU Utilization78.00% %documentation74.90% %-4.0%
    96%
    Throughput875.12 req/sowned875.12 req/s+0.0%
    100%
    Error Rate0.00% %owned0.00% %+0.0%
    100%
    Cost / Hour$0.4545 $/hrprice list$0.4523 $/hr-0.5%
    100%
    Burst — Owned cwm-bench diagnostics
    1,000 RPS target
    holdout measurement
    Engine comparison available
    Traffic burst exceeding normal peak — flash sale, viral event, or coordinated load test.
    Owned metricValueMeaning
    Target RPS1,000target load
    Goodput875.116 RPSwarmup + steady window
    P50 latency1.931 msowned app-local measurement
    P95 latency3.714 msowned app-local measurement
    P99 latency7.017 msowned app-local measurement
    App CPU13.36%owned application host metric
    DB CPU10.91%owned database metric
    DB connections55 / ~500maximum observed / documented ceiling
    CRUD errors≈0%owned canonical CRUD mix
    Cost / hour$0.4545public price list

    Source: company-owned cwm-bench campaign. These are measurements, not simulated values. The scored comparison uses calibrated engine output with owned goodput/error and documentation-backed end-to-end latency/blended-CPU references.

    Methodology & Reference Sources
    How reference values are derived and what each data source covers.

    Pricing accuracy is validated continuously via automated drift checks in CI. See the Simulation Fidelity page for per-provider cost benchmark details across all five cloud providers.

    Want to benchmark a different architecture? Open the Workspace to build and simulate any topology.

    Try the Simulation Yourself

    Load the same 3-tier AWS scenario in the interactive workspace and compare what you see with the reference values on this page.