OpenAI just rolled out CODEX, a new AI coding model that merges its best programming and reasoning capabilities into one faster package while also serving as a key tool in its own training and deployment process
Early versions of CodeX were used to find bugs in its own training runs manage its rollout and analyze evaluation results
CODEX tops agentic coding benchmarks like SWE Bench Pro and Terminal Bench outperforming Opus 4.6 by 12 percent on the latter just minutes after its release
On OS World a benchmark testing AI control of desktop computers the model scored 64.7 percent nearly double the 38.2 percent from the prior Codex version
OpenAI flagged the model as its first high cybersecurity risk rating and committed ten million dollars in API credits to fund defensive security research
It's time we wake up and learn how these new AI models are hacking the future of work