BenchGen is a benchmarking and training platform for AI agents that runs them inside digital-twin simulated environments to capture full decision trajectories and verify real-world performance. It converts benchmark runs into reinforcement learning training data compatible with PPO, GRPO, and PRM methods, enabling continuous agent improvement. BenchGen supports air-gapped and sovereign deployments for government, defense, and regulated industries. The platform serves over 500 teams, has evaluated more than 1,400 agents, and has captured over 2 million trajectories across 750-plus RL environments.