Cloud World Model
RL environments for autoscaling agents
A grounded guide to training autoscaling policies with CWM's RL environment API, valid action structure and evidence limits.
Definition: observe, act, reward, repeat
An RL autoscaling environment lets an agent select a capacity action, receive a new simulated observation and reward, and repeat until an episode ends. In CWM an environment is linked to a simulation; the policy operates on modeled behavior, not live cluster control. Reward shaping can encourage good cost and latency tradeoffs but does not prove a policy will be safe against real traffic.
What the available evidence says
The environment reference documents observation and reward fields and a step/reset loop. The owned AWS accuracy report measures an AWS reference configuration and explicitly discusses which outputs remain documentation-backed; it is not an RL generalization benchmark. Evaluate a trained policy on held-out workload patterns and independent telemetry before considering deployment.
Worked API example: one safe autoscaling step
First create a simulation using the linked walkthrough and save its returned ID. Then create an environment referencing that ID. The step action is a nested object, not a plain string: type no_op with empty parameters is a valid baseline action under the RL request schema. Replace the key and IDs with actual values from your own requests; reset to start another episode.
curl -sS -X POST https://www.cloudworldmodel.ai/api/rl/environments \
-H "Authorization: Bearer $CWM_API_KEY" \
-H "Content-Type: application/json" \
-d "{\"simulationId\":\"$SIM_ID\",\"episodeConfig\":{}}"
# Save the returned environment.id as ENV_ID.
curl -sS -X POST "https://www.cloudworldmodel.ai/api/rl/environments/$ENV_ID/step" \
-H "Authorization: Bearer $CWM_API_KEY" \
-H "Content-Type: application/json" \
-d '{"action":{"type":"no_op","parameters":{}},"tick_seconds":60}'
curl -sS -X POST "https://www.cloudworldmodel.ai/api/rl/environments/$ENV_ID/reset" \
-H "Authorization: Bearer $CWM_API_KEY" \
-H "Content-Type: application/json" -d '{}'Evaluate transfer, not just reward
Separate training scenarios from evaluation scenarios. Record latency, error rates, cost and scaling decisions for baseline and learned policies on identical workloads; compare any proposed live rollout with telemetry and rollback controls. A higher simulator reward only establishes performance against that reward function and model, not production reliability.
