Five breakthroughs that make OpenAI’s o3 a turning point for AI

Stay Ahead, Stay ONMINE

Five breakthroughs that make OpenAI’s o3 a turning point for AI — and one big challenge

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More The end of the year 2024 has brought reckonings for artificial intelligence, as industry insiders feared progress toward even more intelligent AI is slowing down. But OpenAI’s o3 model, announced just last week, has sparked a […]

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More

The end of the year 2024 has brought reckonings for artificial intelligence, as industry insiders feared progress toward even more intelligent AI is slowing down. But OpenAI’s o3 model, announced just last week, has sparked a fresh wave of excitement and debate, and suggests big improvements are still to come in 2025 and beyond.

This model, announced for safety testing among researchers, but not yet released publicly, achieved an impressive score on the important ARC metric. The benchmark was created by François Chollet, a renowned AI researcher and creator of the Keras deep learning framework, and is specifically designed to measure a model’s ability to handle novel, intelligent tasks. As such, it provides a meaningful gauge of progress toward truly intelligent AI systems.

Notably, o3 scored 75.7% on the ARC benchmark under standard compute conditions and 87.5% using high compute, significantly surpassing previous state-of-the-art results, such as the 53% scored by Claude 3.5.

This achievement by o3 represents a surprising advancement, according to Chollet, who had been a critic of the ability of large language models (LLMs) to achieve this sort of intelligence. It highlights innovations that could accelerate progress toward superior intelligence, whether we call it artificial general intelligence (AGI) or not.

AGI is a hyped term, and ill-defined, but it signals a goal: intelligence capable of adapting to novel challenges or questions in ways that surpass human abilities.

OpenAI’s o3 tackles specific hurdles in reasoning and adaptability that have long stymied large language models. At the same time, it exposes challenges, including the high costs and efficiency bottlenecks inherent in pushing these systems to their limits. This article will explore five key innovations behind the o3 model, many of which are underpinned by advancements in reinforcement learning (RL). It will draw on insights from industry leaders, OpenAI’s claims, and above all Chollet’s important analysis, to unpack what this breakthrough means for the future of AI as we move into 2025.

The five core innovations of o3

1. “Program synthesis” for task adaptation

OpenAI’s o3 model introduces a new capability called “program synthesis,” which enables it to dynamically combine things that it learned during pre-training—specific patterns, algorithms, or methods—into new configurations. These things might include mathematical operations, code snippets, or logical procedures that the model has encountered and generalized during its extensive training on diverse datasets. Most significantly, program synthesis allows o3 to address tasks it has never directly seen in training, such as solving advanced coding challenges or tackling novel logic puzzles that require reasoning beyond rote application of learned information. François Chollet describes program synthesis as a system’s ability to recombine known tools in innovative ways—like a chef crafting a unique dish using familiar ingredients. This feature marks a departure from earlier models, which primarily retrieve and apply pre-learned knowledge without reconfiguration — and it’s also one that Chollet had advocated for months ago as the only viable way forward to better intelligence.

2. Natural language program search

At the heart of o3’s adaptability is its use of Chains of Thought (CoTs) and a sophisticated search process that takes place during inference—when the model is actively generating answers in a real-world or deployed setting. These CoTs are step-by-step natural language instructions the model generates to explore solutions. Guided by an evaluator model, o3 actively generates multiple solution paths and evaluates them to determine the most promising option. This approach mirrors human problem-solving, where we brainstorm different methods before choosing the best fit. For example, in mathematical reasoning tasks, o3 generates and evaluates alternative strategies to arrive at accurate solutions. Competitors like Anthropic and Google have experimented with similar approaches, but OpenAI’s implementation sets a new standard.

3. Evaluator model: A new kind of reasoning

O3 actively generates multiple solution paths during inference, evaluating each with the help of an integrated evaluator model to determine the most promising option. By training the evaluator on expert-labeled data, OpenAI ensures that o3 develops a strong capacity to reason through complex, multi-step problems. This feature enables the model to act as a judge of its own reasoning, moving large language models closer to being able to “think” rather than simply respond.

4. Executing Its own programs

One of the most groundbreaking features of o3 is its ability to execute its own Chains of Thought (CoTs) as tools for adaptive problem-solving. Traditionally, CoTs have been used as step-by-step reasoning frameworks to solve specific problems. OpenAI’s o3 extends this concept by leveraging CoTs as reusable building blocks, allowing the model to approach novel challenges with greater adaptability. Over time, these CoTs become structured records of problem-solving strategies, akin to how humans document and refine their learning through experience. This ability demonstrates how o3 is pushing the frontier in adaptive reasoning. According to OpenAI engineer Nat McAleese, o3’s performance on unseen programming challenges, such as achieving a CodeForces rating above 2700, showcases its innovative use of CoTs to rival top competitive programmers. This 2700 rating places the model at “Grandmaster” level, among the top echelon of competitive programmers globally.

5. Deep learning-guided program search

O3 leverages a deep learning-driven approach during inference to evaluate and refine potential solutions to complex problems. This process involves generating multiple solution paths and using patterns learned during training to assess their viability. François Chollet and other experts have noted that this reliance on ‘indirect evaluations’—where solutions are judged based on internal metrics rather than tested in real-world scenarios—can limit the model’s robustness when applied to unpredictable or enterprise-specific contexts.

Additionally, o3’s dependence on expert-labeled datasets for training its evaluator model raises concerns about scalability. While these datasets enhance precision, they also require significant human oversight, which can restrict the system’s adaptability and cost-efficiency. Chollet highlights that these trade-offs illustrate the challenges of scaling reasoning systems beyond controlled benchmarks like ARC-AGI.

Ultimately, this approach demonstrates both the potential and limitations of integrating deep learning techniques with programmatic problem-solving. While o3’s innovations showcase progress, they also underscore the complexities of building truly generalizable AI systems.

The big challenge to o3

OpenAI’s o3 model achieves impressive results but at significant computational cost, consuming millions of tokens per task — and this costly approach is model’s biggest challenge. François Chollet, Nat McAleese, and others highlight concerns about the economic feasibility of such models, emphasizing the need for innovations that balance performance with affordability.

The o3 release has sparked attention across the AI community. Competitors such as Google with Gemini 2 and Chinese firms like DeepSeek 3 are also advancing, making direct comparisons challenging until these models are more widely tested.

Opinions on o3 are divided: some laud its technical strides, while others cite high costs and a lack of transparency, suggesting its real value will only become clear with broader testing. One of the biggest critiques came from Google DeepMind’s Denny Zhou, who implicitly attacked the model’s reliance on reinforcement learning (RL) scaling and search mechanisms as a potential “dead end,” arguing instead that a model should be able to learn to reason from simpler fine-tuning processes.

What this means for enterprise AI

Whether or not it represents the perfect direction for further innovation, for enterprises, o3’s new-found adaptability shows that AI will in one way or another continue to transform industries, from customer service and scientific research, in the future.

Industry players will need some time to digest what o3 has delivered here. For enterprises concerned about o3’s high computational costs, OpenAI’s upcoming release of the scaled-down “o3-mini” version of the model provides a potential alternative. While it sacrifices some of the full model’s capabilities, o3-mini promises a more affordable option for businesses to experiment with — retaining much of the core innovation while significantly reducing test-time compute requirements.

It may be some time before enterprise companies can get their hands on the o3 model. OpenAI says the o3-mini is expected to launch by the end of January. The full o3 release will follow after, though the timelines depend on feedback and insights gained during the current safety testing phase. Enterprise companies will be well advised to test it out. They’ll want to ground the model with their data and use cases and see how it really works.

But in the mean time, they can already use the many other competent models that are already out and well tested, including the flagship o4 model and other competing models — many of which are already robust enough for building intelligent, tailored applications that deliver practical value.

Indeed, next year, we’ll be operating on two gears. The first is in achieving practical value from AI applications, and fleshing out what models can do with AI agents, and other innovations already achieved. The second will be sitting back with the popcorn and seeing how the intelligence race plays out — and any progress will just be icing on the cake that has already been delivered.

For more on o3’s innovations, watch the full YouTube discussion between myself and Sam Witteveen below, and follow VentureBeat for ongoing coverage of AI advancements.

Daily insights on business use cases with VB Daily

If you want to impress your boss, VB Daily has you covered. We give you the inside scoop on what companies are doing with generative AI, from regulatory shifts to practical deployments, so you can share insights for maximum ROI.

Read our Privacy Policy

Thanks for subscribing. Check out more VB newsletters here.

An error occured.

Stay Ahead

Explore More Insights

Stay ahead with more perspectives on cutting-edge power, infrastructure, energy, bitcoin and AI solutions. Explore these articles to uncover strategies and insights shaping the future of industries.

Oncor to propose rate increase, says 5-year capital expenditures could ‘soar’ by $12B

Dive Brief: Oncor Electric said Thursday that it will ask Texas regulators to increase its base rates by approximately $834 million to support Texas investments that include the utility’s $36 billion five-year capital spending plan — though that spending figure is likely to climb, officials said. In February, Oncor announced a

Chronosphere unveils logging package with cost control features

According to a study by Chronosphere, enterprise log data is growing at 250% year-over-year, and Chronosphere Logs helps engineers and observability teams to resolve incidents faster while controlling costs. The usage and volume analysis and proactive recommendations can help reduce data before it’s stored, the company says. “Organizations are drowning

Cisco CIO on the future of IT: AI, simplicity, and employee power

AI can democratize access to information to deliver a “white-glove experience” once reserved for senior executives, Previn said. That might include, for example, real-time information retrieval and intelligent process execution for every employee. “Usually, in a large company, you’ve got senior executives, and you’ve got early career hires, and it’s

AMI MegaRAC authentication bypass flaw is being exploitated, CISA warns

The spoofing attack works by manipulating HTTP request headers sent to the Redfish interface. Attackers can add specific values to headers like “X-Server-Addr” to make their external requests appear as if they’re coming from inside the server itself. Since the system automatically trusts internal requests as authenticated, this spoofing technique

Israeli Gas Flows to Egypt Return to Normal as Iran Truce Holds

Israeli natural gas flows to Egypt returned to normal levels after a truce with Iran allowed the Jewish state to reopen facilities shuttered by the 12-day conflict. Daily exports have climbed to 1 billion cubic feet per day, according to two people with direct knowledge of the situation. That’s up from 260 million cubic feet when Israel’s Leviathan gas field, the country’s biggest, restarted on Wednesday, they said, declining to be identified because they’re not authorized to speak to the media. The increased flows have let Egyptian authorities resume supplies to some factories that had been halted because of the shortages. Israel temporarily closed two of its three gas fields – Chevron-operated Leviathan and Energean’s Karish – shortly after launching attacks on Iran on June 13. The facilities that provided the bulk of exports to Egypt and Jordan resumed operations last week after a US-brokered ceasefire with the Islamic Republic took hold. The ramped-up supplies are a relief for Cairo, which has swung from a net exporter to importer of natural gas in recent years. As Israel and Iran traded blows, Egypt enacted contingency plans that included seeking alternative fuel purchases, limiting gas to some industries and switching power stations to fuel oil and diesel to maintain electricity output. What do you think? We’d love to hear from you, join the conversation on the Rigzone Energy Network. The Rigzone Energy Network is a new social experience created for you and all energy professionals to Speak Up about our industry, share knowledge, connect with peers and industry insiders and engage in a professional community that will empower your career in energy.

California Regulator Wants to Pause Newsom Refinery Profit Cap

California’s energy market regulator is backing off a plan to place a profit cap on oil refiners in the state. Siva Gunda, vice chair of the California Energy Commission, said during a Friday briefing that the cap would “serve as a deterrent” to refiners boosting investments in the state. Gunda said the commission wants to increase gasoline supply in California after two refineries announced plans to close in the next year, accounting for about one-fifth of the state’s crude-processing capacity. The recommendation marks a reversal from years of regulatory scrutiny by Governor Gavin Newsom and the California Energy Commission that contributed to plans by Phillips 66 and Valero Energy Corp. to shut their refineries. The closings prompted Newsom to adjust course in April and urge the energy regulator to collaborate with fuel makers to ensure affordable and reliable supply. Gunda wrote in a Friday letter to Newsom that the commission should pause implementation of a profit margin cap and focus on fuel resupply strategies instead. It comes more than two years after Newsom and state lawmakers gave the energy commission authority to determine a profit margin on refiners and impose financial penalties for violations. The state will be looking to increase fuel imports to make up for the loss of refining capacity, Gunda said. In the short term, California gas prices could rise 15 to 30 cents a gallon because of the loss of production, he said. A spokesperson for the energy commission said the estimated price increases would be mitigated by the plan presented on Friday. Californians already pay the highest gasoline prices in the country. Wade Crowfoot, secretary of the California Natural Resources Agency, said residents want the state to transition away from oil and gas yet they need to prevent cost spikes. “We get it,” he said.

State utility regulators urge FERC to slash ROE transmission incentive

Utility regulators from about 35 states are urging the Federal Energy Regulatory Commission to sharply limit a 0.5% return on equity incentive the agency gives to utilities that join regional transmission organizations. “The time has come for the Commission to eliminate its policy of granting the RTO Participation Adder in perpetuity, if not to eliminate this incentive altogether,” the Organization of PJM States, the Organization of MISO States, the New England States Committee on Electricity and the Southwest Power Pool Regional State Committee said in a Friday letter to FERC. The state regulators and others contends the RTO incentive adds millions to ratepayer costs to encourage behavior — being an RTO member — that they would likely do anyway. In 2021, FERC proposed limiting its ROE adder to three years. FERC Chairman Mark Christie supports the proposal as well as limiting other incentives aimed at encouraging utilities to build transmission lines. However, it appears he has been unable to convince a majority of FERC commissioners to reduce those incentives. Christie’s term ends today, although he plans to stay at the agency until at least FERC’s next open meeting on July 24. Limiting the ROE incentive could reduce utility income. Public Service Enterprise Group, for example, estimates that removing the incentive would cut annual net income and cash inflows by about $40 million for its Public Service Electric & Gas subsidiary, according to a Feb. 25 filing at the U.S. Securities and Exchange Commission. The utility earned about $1.5 billion in 2024. Ending the incentive would reduce American Electric Power’s pretax income by $35 million to $50 million a year, the utility company said in its 2023 annual report with the SEC. In April, WIRES, a transmission-focused trade group, the Edison Electric Institute, which represents investor-owned utilities, and GridWise Alliance, a

Affordability a ‘formidable challenge’ as load shifts to tech, industrial customers: ICF

Dive Brief: Keeping electricity affordable for consumers is a “formidable challenge” amid projections of declining generation capacity reserves and persistent uncertainty around the scale and pace of future load growth, ICF International Vice President of Energy Markets Maria Scheller said Thursday. Meanwhile, broad policy uncertainty and an increasingly shaky regulatory environment give utilities and capital markets pause about expensive new infrastructure investments that could become stranded assets, Scheller said in a webinar on ICF’s “Powering the Future: Addressing Surging U.S. Electricity Demand” report. Policy conversations around import tariffs, federal energy tax credits and permitting reform are unfolding as the balance of electricity demand shifts from residential and business consumers to technology and industrial customers, which tend to require around-the-clock power, Scheller added. Dive Insight: The coming shift in U.S. electricity consumption represents less of a new paradigm than a return to the industrial-driven demand the country saw from the 1950s into the 1980s, after which deindustrialization and consumer-centric trends like the widespread adoption of air conditioning, electric resistance heating and personal computing shifted the balance toward the residential segment, Scheller said. The shift is important because unlike residential loads, which show considerable seasonal and intraday variation, industrial loads are flatter, less weather-dependent and more sensitive to voltage fluctuations, Scheller said. By 2035, ICF expects nearly 40% of total U.S. load will have a “flat, power-quality-sensitive profile,” and that overall load will grow faster than peak load, she said. In 2030, ICF projects more than 3% annual power consumption growth, compared with less than 2% annual peak load growth, according to a webinar slide. That’s not to say residential demand won’t also grow in the next few years as consumers electrify home heating and buy more electric vehicles — only that data centers and other industrial demand will “dwarf” it, Scheller

Trump attacks on NRC independence pose health, safety risks

Edwin Lyman is director of nuclear power safety at the Union of Concerned Scientists. A White House executive order issued last month targeting the independence of the Nuclear Regulatory Commission, the federal agency that oversees the safety and security of U.S. commercial nuclear facilities and materials, as well as the possibly illegal firing earlier this month of Commissioner Christopher Hanson by President Donald Trump, are raising serious concerns about the agency’s effectiveness as a regulator going forward. While I’ve often been a critic of the NRC for taking actions favoring the nuclear industry at the expense of public health and safety, preserving the NRC in its current form is the best hope for heading off a U.S. nuclear plant disaster like the 2011 Fukushima Daiichi reactor meltdowns in Japan. My long-standing beef with the NRC has primarily been with its political leadership, not with the rank-and-file staff of highly knowledgeable inspectors, analysts and researchers committed to helping ensure that nuclear power remains safe and secure. These professionals are well aware how quickly things can go south at a nuclear power plant without rigorous oversight. They know from experience what obscure corners to look in and what questions to ask. And they can tell — and are not afraid to push back — when they are getting sold snake oil by fly-by-night startups looking to make easy money by capitalizing on the current nuclear power craze. Technical rigor and expert judgment form the bedrock of this work. But when staff are compelled to sweep legitimate safety concerns under the rug in the interest of political expediency, many will leave rather than compromise their scientific integrity. So there is little wonder that a wave of experienced personnel is headed out the door in the wake of the executive order on NRC “reform,”

North America Loses Rigs Week on Week

North America dropped six rigs week on week, according to Baker Hughes’ latest North America rotary rig count, which was released on June 27. Although the U.S. dropped seven rigs week on week, Canada added one rig during the same timeframe, taking the total North America rig count down to 687, comprising 547 rigs from the U.S. and 140 rigs from Canada, the count outlined. Of the total U.S. rig count of 547, 533 rigs are categorized as land rigs, 12 are categorized as offshore rigs, and two are categorized as inland water rigs. The total U.S. rig count is made up of 432 oil rigs, 109 gas rigs, and six miscellaneous rigs, according to Baker Hughes’ count, which revealed that the U.S. total comprises 496 horizontal rigs, 38 directional rigs, and 13 vertical rigs. Week on week, the U.S. land rig count reduced by five, its offshore rig count decreased by two, and its inland water rig count remained unchanged, the count highlighted. The country’s oil rig count dropped by six, its gas rig count dropped by two, and its miscellaneous rig count increased by one, week on week, the count showed. The U.S. horizontal rig count dropped by six, its directional rig count dropped by two, and its vertical rig count increased by one, week on week, the count revealed. A major state variances subcategory included in the rig count showed that, week on week, Wyoming dropped five rigs, and Oklahoma, Louisiana, and Colorado each dropped one rig. A major basin variances subcategory included in Baker Hughes’ rig count showed that, week on week, the Granite Wash basin dropped one rig and the Permian basin dropped one rig. Canada’s total rig count of 140 is made up of 94 oil rigs and 46 gas rigs, Baker Hughes pointed

Datacenter industry calls for investment after EU issues water consumption warning

CISPE’s response to the European Commission’s report warns that the resulting regulatory uncertainty could hurt the region’s economy. “Imposing new, standalone water regulations could increase costs, create regulatory fragmentation, and deter investment. This risks shifting infrastructure outside the EU, undermining both sustainability and sovereignty goals,” CISPE said in its latest policy recommendation, Advancing water resilience through digital innovation and responsible stewardship. “Such regulatory uncertainty could also reduce Europe’s attractiveness for climate-neutral infrastructure investment at a time when other regions offer clear and stable frameworks for green data growth,” it added. CISPE’s recommendations are a mix of regulatory harmonization, increased investment, and technological improvement. Currently, water reuse regulation is directed towards agriculture. Updated regulation across the bloc would encourage more efficient use of water in industrial settings such as datacenters, the asosciation said. At the same time, countries struggling with limited public sector budgets are not investing enough in water infrastructure. This could only be addressed by tapping new investment by encouraging formal public-private partnerships (PPPs), it suggested: “Such a framework would enable the development of sustainable financing models that harness private sector innovation and capital, while ensuring robust public oversight and accountability.” Nevertheless, better water management would also require real-time data gathered through networks of IoT sensors coupled to AI analytics and prediction systems. To that end, cloud datacenters were less a drain on water resources than part of the answer: “A cloud-based approach would allow water utilities and industrial users to centralize data collection, automate operational processes, and leverage machine learning algorithms for improved decision-making,” argued CISPE.

HPE-Juniper deal clears DOJ hurdle, but settlement requires divestitures

In HPE’s press release following the court’s decision, the vendor wrote that “After close, HPE will facilitate limited access to Juniper’s advanced Mist AIOps technology.” In addition, the DOJ stated that the settlement requires HPE to divest its Instant On business and mandates that the merged firm license critical Juniper software to independent competitors. Specifically, HPE must divest its global Instant On campus and branch WLAN business, including all assets, intellectual property, R&D personnel, and customer relationships, to a DOJ-approved buyer within 180 days. Instant On is aimed primarily at the SMB arena and offers a cloud-based package of wired and wireless networking gear that’s designed for so-called out-of-the-box installation and minimal IT involvement, according to HPE. HPE and Juniper focused on the positive in reacting to the settlement. “Our agreement with the DOJ paves the way to close HPE’s acquisition of Juniper Networks and preserves the intended benefits of this deal for our customers and shareholders, while creating greater competition in the global networking market,” HPE CEO Antonio Neri said in a statement. “For the first time, customers will now have a modern network architecture alternative that can best support the demands of AI workloads. The combination of HPE Aruba Networking and Juniper Networks will provide customers with a comprehensive portfolio of secure, AI-native networking solutions, and accelerate HPE’s ability to grow in the AI data center, service provider and cloud segments.” “This marks an exciting step forward in delivering on a critical customer need – a complete portfolio of modern, secure networking solutions to connect their organizations and provide essential foundations for hybrid cloud and AI,” said Juniper Networks CEO Rami Rahim. “We look forward to closing this transaction and turning our shared vision into reality for enterprise, service provider and cloud customers.”

Data center costs surge up to 18% as enterprises face two-year capacity drought

“AI workloads, especially training and archival, can absorb 10-20ms latency variance if offset by 30-40% cost savings and assured uptime,” said Gogia. “Des Moines and Richmond offer better interconnection diversity today than some saturated Tier-1 hubs.” Contract flexibility is also crucial. Rather than traditional long-term leases, enterprises are negotiating shorter agreements with renewal options and exploring revenue-sharing arrangements tied to business performance. Maximizing what you have With expansion becoming more costly, enterprises are getting serious about efficiency through aggressive server consolidation, sophisticated virtualization and AI-driven optimization tools that squeeze more performance from existing space. The companies performing best in this constrained market are focusing on optimization rather than expansion. Some embrace hybrid strategies blending existing on-premises infrastructure with strategic cloud partnerships, reducing dependence on traditional colocation while maintaining control over critical workloads. The long wait When might relief arrive? CBRE’s analysis shows primary markets had a record 6,350 MW under construction at year-end 2024, more than double 2023 levels. However, power capacity constraints are forcing aggressive pre-leasing and extending construction timelines to 2027 and beyond. The implications for enterprises are stark: with construction timelines extending years due to power constraints, companies are essentially locked into current infrastructure for at least the next few years. Those adapting their strategies now will be better positioned when capacity eventually returns.

Cisco backs quantum networking startup Qunnect

In partnership with Deutsche Telekom’s T-Labs, Qunnect has set up quantum networking testbeds in New York City and Berlin. “Qunnect understands that quantum networking has to work in the real world, not just in pristine lab conditions,” Vijoy Pandey, general manager and senior vice president of Outshift by Cisco, stated in a blog about the investment. “Their room-temperature approach aligns with our quantum data center vision.” Cisco recently announced it is developing a quantum entanglement chip that could ultimately become part of the gear that will populate future quantum data centers. The chip operates at room temperature, uses minimal power, and functions using existing telecom frequencies, according to Pandey.

HPE announces GreenLake Intelligence, goes all-in with agentic AI

Like a teammate who never sleeps Agentic AI is coming to Aruba Central as well, with an autonomous supervisory module talking to multiple specialized models to, for example, determine the root cause of an issue and provide recommendations. David Hughes, SVP and chief product officer, HPE Aruba Networking, said, “It’s like having a teammate who can work while you’re asleep, work on problems, and when you arrive in the morning, have those proposed answers there, complete with chain of thought logic explaining how they got to their conclusions.” Several new services for FinOps and sustainability in GreenLake Cloud are also being integrated into GreenLake Intelligence, including a new workload and capacity optimizer, extended consumption analytics to help organizations control costs, and predictive sustainability forecasting and a managed service mode in the HPE Sustainability Insight Center. In addition, updates to the OpsRamp operations copilot, launched in 2024, will enable agentic automation including conversational product help, an agentic command center that enables AI/ML-based alerts, incident management, and root cause analysis across the infrastructure when it is released in the fourth quarter of 2025. It is now a validated observability solution for the Nvidia Enterprise AI Factory. OpsRamp will also be part of the new HPE CloudOps software suite, available in the fourth quarter, which will include HPE Morpheus Enterprise and HPE Zerto. HPE said the new suite will provide automation, orchestration, governance, data mobility, data protection, and cyber resilience for multivendor, multi cloud, multi-workload infrastructures. Matt Kimball, principal analyst for datacenter, compute, and storage at Moor Insights & strategy, sees HPE’s latest announcements aligning nicely with enterprise IT modernization efforts, using AI to optimize performance. “GreenLake Intelligence is really where all of this comes together. I am a huge fan of Morpheus in delivering an agnostic orchestration plane, regardless of operating stack

MEF goes beyond metro Ethernet, rebrands as Mplify with expanded scope on NaaS and AI

While MEF is only now rebranding, Vachon said that the scope of the organization had already changed by 2005. Instead of just looking at metro Ethernet, the organization at the time had expanded into carrier Ethernet requirements. The organization has also had a growing focus on solving the challenge of cross-provider automation, which is where the LSO framework fits in. LSO provides the foundation for an automation framework that allows providers to more efficiently deliver complex services across partner networks, essentially creating a standardized language for service integration. NaaS leadership and industry blueprint Building on the LSO automation framework, the organization has been working on efforts to help providers with network-as-a-service (NaaS) related guidance and specifications. The organization’s evolution toward NaaS reflects member-driven demands for modern service delivery models. Vachon noted that MEF member organizations were asking for help with NaaS, looking for direction on establishing common definitions and some standard work. The organization responded by developing comprehensive industry guidance. “In 2023 we launched the first blueprint, which is like an industry North Star document. It includes what we think about NaaS and the work we’re doing around it,” Vachon said. The NaaS blueprint encompasses the complete service delivery ecosystem, with APIs including last mile, cloud, data center and security services. (Read more about its vision for NaaS, including easy provisioning and integrated security across a federated network of providers)

Microsoft will invest $80B in AI data centers in fiscal 2025

And Microsoft isn’t the only one that is ramping up its investments into AI-enabled data centers. Rival cloud service providers are all investing in either upgrading or opening new data centers to capture a larger chunk of business from developers and users of large language models (LLMs). In a report published in October 2024, Bloomberg Intelligence estimated that demand for generative AI would push Microsoft, AWS, Google, Oracle, Meta, and Apple would between them devote $200 billion to capex in 2025, up from $110 billion in 2023. Microsoft is one of the biggest spenders, followed closely by Google and AWS, Bloomberg Intelligence said. Its estimate of Microsoft’s capital spending on AI, at $62.4 billion for calendar 2025, is lower than Smith’s claim that the company will invest $80 billion in the fiscal year to June 30, 2025. Both figures, though, are way higher than Microsoft’s 2020 capital expenditure of “just” $17.6 billion. The majority of the increased spending is tied to cloud services and the expansion of AI infrastructure needed to provide compute capacity for OpenAI workloads. Separately, last October Amazon CEO Andy Jassy said his company planned total capex spend of $75 billion in 2024 and even more in 2025, with much of it going to AWS, its cloud computing division.

John Deere unveils more autonomous farm machines to address skill labor shortage

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Self-driving tractors might be the path to self-driving cars. John Deere has revealed a new line of autonomous machines and tech across agriculture, construction and commercial landscaping. The Moline, Illinois-based John Deere has been in business for 187 years, yet it’s been a regular as a non-tech company showing off technology at the big tech trade show in Las Vegas and is back at CES 2025 with more autonomous tractors and other vehicles. This is not something we usually cover, but John Deere has a lot of data that is interesting in the big picture of tech. The message from the company is that there aren’t enough skilled farm laborers to do the work that its customers need. It’s been a challenge for most of the last two decades, said Jahmy Hindman, CTO at John Deere, in a briefing. Much of the tech will come this fall and after that. He noted that the average farmer in the U.S. is over 58 and works 12 to 18 hours a day to grow food for us. And he said the American Farm Bureau Federation estimates there are roughly 2.4 million farm jobs that need to be filled annually; and the agricultural work force continues to shrink. (This is my hint to the anti-immigration crowd). John Deere’s autonomous 9RX Tractor. Farmers can oversee it using an app. While each of these industries experiences their own set of challenges, a commonality across all is skilled labor availability. In construction, about 80% percent of contractors struggle to find skilled labor. And in commercial landscaping, 86% of landscaping business owners can’t find labor to fill open positions, he said. “They have to figure out how to do

2025 playbook for enterprise AI success, from agents to evals

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More 2025 is poised to be a pivotal year for enterprise AI. The past year has seen rapid innovation, and this year will see the same. This has made it more critical than ever to revisit your AI strategy to stay competitive and create value for your customers. From scaling AI agents to optimizing costs, here are the five critical areas enterprises should prioritize for their AI strategy this year. 1. Agents: the next generation of automation AI agents are no longer theoretical. In 2025, they’re indispensable tools for enterprises looking to streamline operations and enhance customer interactions. Unlike traditional software, agents powered by large language models (LLMs) can make nuanced decisions, navigate complex multi-step tasks, and integrate seamlessly with tools and APIs. At the start of 2024, agents were not ready for prime time, making frustrating mistakes like hallucinating URLs. They started getting better as frontier large language models themselves improved. “Let me put it this way,” said Sam Witteveen, cofounder of Red Dragon, a company that develops agents for companies, and that recently reviewed the 48 agents it built last year. “Interestingly, the ones that we built at the start of the year, a lot of those worked way better at the end of the year just because the models got better.” Witteveen shared this in the video podcast we filmed to discuss these five big trends in detail. Models are getting better and hallucinating less, and they’re also being trained to do agentic tasks. Another feature that the model providers are researching is a way to use the LLM as a judge, and as models get cheaper (something we’ll cover below), companies can use three or more models to

OpenAI’s red teaming innovations define new essentials for security leaders in the AI era

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More OpenAI has taken a more aggressive approach to red teaming than its AI competitors, demonstrating its security teams’ advanced capabilities in two areas: multi-step reinforcement and external red teaming. OpenAI recently released two papers that set a new competitive standard for improving the quality, reliability and safety of AI models in these two techniques and more. The first paper, “OpenAI’s Approach to External Red Teaming for AI Models and Systems,” reports that specialized teams outside the company have proven effective in uncovering vulnerabilities that might otherwise have made it into a released model because in-house testing techniques may have missed them. In the second paper, “Diverse and Effective Red Teaming with Auto-Generated Rewards and Multi-Step Reinforcement Learning,” OpenAI introduces an automated framework that relies on iterative reinforcement learning to generate a broad spectrum of novel, wide-ranging attacks. Going all-in on red teaming pays practical, competitive dividends It’s encouraging to see competitive intensity in red teaming growing among AI companies. When Anthropic released its AI red team guidelines in June of last year, it joined AI providers including Google, Microsoft, Nvidia, OpenAI, and even the U.S.’s National Institute of Standards and Technology (NIST), which all had released red teaming frameworks. Investing heavily in red teaming yields tangible benefits for security leaders in any organization. OpenAI’s paper on external red teaming provides a detailed analysis of how the company strives to create specialized external teams that include cybersecurity and subject matter experts. The goal is to see if knowledgeable external teams can defeat models’ security perimeters and find gaps in their security, biases and controls that prompt-based testing couldn’t find. What makes OpenAI’s recent papers noteworthy is how well they define using human-in-the-middle

Stay Ahead, Stay ONMINE