Programmable Worlds
for AI Agents

Sandboxed, observable environments to train, evaluate, and verify intelligent systems.

forge::runtimeSANDBOX ACTIVE
$forge run gmail--agent anthropic:claude-3-7-sonnet --policy grpo
EnvironmentGmail Sandbox (34 emails, 19 contacts, auto-reply)
ExecutionEpisode #412 · Step 14/50 · Trajectory recorded
VerifiersExactStateVerifier · EventVerifier · LayeredVerifier
Verdict✓ All checks passed · Reward: 0.94 · Exported to grpo_rollouts.parquet
Held-Out Pass
72.4%
Reward Hacking
< 0.02%
Throughput
3,850 rollouts/s
CORE DOMAINS

Four Pillars of Programmable Worlds

REINFORCEMENT LEARNING

RL Gym Environments for Agent Training

Train agents on real applications with live state. Collect rollout trajectories, score episodes with layered verifiers, and export directly to policy fine-tuning datasets.

Premade & Custom Apps

Pre-seeded Gmail and Slack replicas, Ubuntu shell, Chromium browser, or custom FastAPI apps.

Computed Ground Truth

Exact state, event sequences, and temporal logic verifiers—no model grading contamination.

Direct Policy Export

Native serialization to GRPO rollouts, DPO preference pairs, and SFT datasets.

forge::env::grpo_rolloutPolicy Loop
# Python SDK
env = forge.make("gmail-v1", agent="anthropic:claude-3-7")
obs, info = env.reset(seed=42)
action = agent.predict(obs)
dataset.export("grpo_rollouts.parquet")
Held-Out Pass
72.4%
Reward Hacking
< 0.02%
Variance
0.08
Speedup
180x
SYSTEM ARCHITECTURE

Engineered for Fidelity

realistic simulation, verified reward signals, and policy training.

⚙️01

Gymnasium Environment Facade

Shared contracts across all environments for reset, state management, tool schemas, and step execution with zero state leaks.

GymnasiumToolProviderObservationEncoderMicroVMs
⚖️02

Layered Verification Suite

Six built-in verifiers compute mathematical verdicts from state transitions and event ordering, enforcing grader independence.

ExactStateTemporalVerifierNegativeVerifierPolicyVerifier
🔁03

Replay & Branch Engine

Deterministically re-run episodes step-by-step or fork from step T to test alternate action paths and edge cases.

Time-TravelForkingDeterministic ClockState Snapshot
📦04

Dataset Exporters

Convert completed trajectories directly into GRPO rollouts, DPO preference pairs, SFT pairs, and failure datasets.

GRPO RolloutsDPO PairsSFT PairsParquet / JSONL
🔍05

Trace & Loss Taxonomy

Correlates prompts, tool calls, and state changes into unified traces with automated 7-mode failure classification.

Unified Traces7-Mode LossReward-Hacking AuditAnomalies
📊06

Generalization Benchmark

Evaluates trained checkpoints strictly on held-out environments to measure pass rates and reward stability across seeds.

Held-Out SplitsZero LeakageReproducible SeedsHarbor Integration

Build Programmable Worlds

Spin up sandboxed environments, evaluate autonomous agents with computed ground truth, and train policies on real-world applications.