OpenAI
1.981
113
OpenAI’s latest features take direct aim at the app store model
OpenAI is building out the pieces of an alternative to the traditional app store model, turning ChatGPT into a place where software can be discovered…
OpenAI reportedly in talks to raise 30B round at 1.4T valuation
The new round is anticipated to be the company's last before its delayed 2027 public debut.
OpenAI repotedly in talks to raise 30B round at 1.4T valuation
The new round is anticipated to be the company's last before its delayed 2027 public debut.
Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock
GPT-6.1 Sol is now generally available on Amazon Bedrock, bringing stronger reasoning to coding, computer use, and professional workloads that run frequently.
Here’s why OpenAI is absent from Nvidia’s industry-wide effort to end rogue AI agents
OpenAI isn't a public supporter of Nvidia's Open Agent Safety Platform, but it is privately working with Nvidia, TechCrunch has learned.
Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.
OpenAI takes on Microsoft with the launch of what feels a whole lot like ChatGPT’s own office suite
OpenAI's newly announced suite of office features puts it into mor direct competition with more traditional software companies.
AI researchers put out videos saying superintelligence is ‘exactly as dangerous as it sounds’
"The chance of human extinction is about a coin flip, in my view," Geoffrey Irving, a former OpenAI and Google DeepMind employee, said in a…
ChatGPT Pro 500
UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed its assigned role. We introduce UserProxyBench, an evaluation layer over the tau-bench family, and the User Fidelity Score (UFS), which measures adherence to the benchmark's private user instructions using task-grounded rubric criteria scored independently of agent success. Holding the agent fixed at GPT-5.5 and varying only the user proxy across 375 enterprise tasks changes mean task reward by 15.2 points, while 24.4% of successful episodes contain a user-specification violation. The dominant failure is premature disclosure: users provide information before it is requested. This behavior has little effect on task reward, yet among successful episodes it causes the agent to make 1.06 fewer tool calls on average, changing the interaction being evaluated while preserving the reward. Finally, across seven proxies we identify an empirical cost-fidelity frontier, enabling practitioners to select the least expensive simulator that satisfies a required fidelity level.
OpenAI launches Dots, its bubbly agentic avatar
Dots are meant to operate independent of any specific hardware or interface, pursuing user-defined goals continuously in the background with minimal oversight.
OpenAI gives Codex reusable cloud environments that work across devices
OpenAI is expanding Codex with reusable cloud development environments, a revamped CLI with voice controls, new code review tools and a security-focused product for scanning…
OpenAI launches GPT-6.1 Sol, says it nearly matches GPT-6 Astra and costs less
OpenAI says GPT-6.1 Sol delivers significant improvements over GPT-6 Sol across complex professional tasks, including code writing and debugging, document understanding, and executing multi-step business…
OpenAI expands ChatGPT’s plugins with app-like interfaces and automations
OpenAI is expanding ChatGPT plugins with dedicated sidebar homes, interactive panels, file viewers, improved discovery, and support for automations.
OpenAI launches Dots, its Muse competitor
OpenAI is responding to Meta's buzzy Muse AI with agentic helpers of its own: Dots. During its DevDay keynote on Tuesday, OpenAI announced that Dots…