Zro is a private inference endpoint for open-weight coding models, served from EU infrastructure with zero request retention and no training on customer data. It supports popular coding agents including Claude Code, Codex CLI, Cursor, Cline, and others via an npm package and OpenAI- or Anthropic-compatible APIs. Zro is optimized for long-context, multi-turn coding sessions using HyperQuant compression and custom attention kernels on AMD, NVIDIA, and TPU hardware. Current available models include MiniMax M3 and GLM-5.2, with additional open-weight models planned. Billing starts at $20 per month for $60 of inference spend, with pay-as-you-go usage packs also available.