Cloud World Model on Smithery

    Cloud World Model

    RL environments for autoscaling agents

    A grounded guide to training autoscaling policies with CWM's RL environment API, valid action structure and evidence limits.

    Definition: observe, act, reward, repeat

    An RL autoscaling environment lets an agent select a capacity action, receive a new simulated observation and reward, and repeat until an episode ends. In CWM an environment is linked to a simulation; the policy operates on modeled behavior, not live cluster control. Reward shaping can encourage good cost and latency tradeoffs but does not prove a policy will be safe against real traffic.

    RL environment documentation →

    What the available evidence says

    The environment reference documents observation and reward fields and a step/reset loop. The owned AWS accuracy report measures an AWS reference configuration and explicitly discusses which outputs remain documentation-backed; it is not an RL generalization benchmark. Evaluate a trained policy on held-out workload patterns and independent telemetry before considering deployment.

    Read the accuracy benchmark →

    Worked API example: one safe autoscaling step

    First create a simulation using the linked walkthrough and save its returned ID. Then create an environment referencing that ID. The step action is a nested object, not a plain string: type no_op with empty parameters is a valid baseline action under the RL request schema. Replace the key and IDs with actual values from your own requests; reset to start another episode.

    curl -sS -X POST https://www.cloudworldmodel.ai/api/rl/environments \
      -H "Authorization: Bearer $CWM_API_KEY" \
      -H "Content-Type: application/json" \
      -d "{\"simulationId\":\"$SIM_ID\",\"episodeConfig\":{}}"
    # Save the returned environment.id as ENV_ID.
    curl -sS -X POST "https://www.cloudworldmodel.ai/api/rl/environments/$ENV_ID/step" \
      -H "Authorization: Bearer $CWM_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{"action":{"type":"no_op","parameters":{}},"tick_seconds":60}'
    curl -sS -X POST "https://www.cloudworldmodel.ai/api/rl/environments/$ENV_ID/reset" \
      -H "Authorization: Bearer $CWM_API_KEY" \
      -H "Content-Type: application/json" -d '{}'

    Read RL action and episode reference →

    Evaluate transfer, not just reward

    Separate training scenarios from evaluation scenarios. Record latency, error rates, cost and scaling decisions for baseline and learned policies on identical workloads; compare any proposed live rollout with telemetry and rollback controls. A higher simulator reward only establishes performance against that reward function and model, not production reliability.

    Model fidelity and caveats →