Stay Ahead, Stay ONMINE

The rise of browser-use agents: Why Convergence’s Proxy is beating OpenAI’s Operator

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More A new wave of AI-powered browser-use agents is emerging, promising to transform how enterprises interact with the web. These agents can autonomously navigate websites, retrieve information, and even complete transactions – but early testing reveals significant […]

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More


A new wave of AI-powered browser-use agents is emerging, promising to transform how enterprises interact with the web. These agents can autonomously navigate websites, retrieve information, and even complete transactions – but early testing reveals significant gaps between promise and performance.

While consumer examples offered by OpenAI’s new browser-use agent Operator, like ordering pizza or buying game tickets, have grabbed headlines, the question is about where the main developer and enterprise use cases are. “The thing that we don’t know is what will be the killer app,” said Sam Witteveen, co-founder of Red Dragon, a company that develops AI agent applications. “My guess is it’s going to be things that just take time on the web that you don’t actually enjoy.” This includes things like going on the web and searching for the cheapest price of a product or booking the best hotel accommodations. More likely it will be used in combination with other tools like Deep Research, where companies can then do even more sophisticated research plus execution of tasks around the web.

Companies need to carefully evaluate the rapidly evolving landscape as established players and startups take different approaches to solving the autonomous browsing challenge.

Key players in the browser-use agent landscape

The field has quickly become crowded with both major tech companies and innovative startups:

Operator and Proxy are the most advanced, in terms of being consumer-friendly and out-of-the-box ready. Many of the others appear to be positioning themselves more for developer or enterprise usage. For example, Browser Use, a Y-Combinator startup that allows users to customize the models used with the agent. This gives you more control over how the agent works, including using a model from your local machine. But it’s definitely more involved.

The others listed above provide a varying degree of functionality and interaction with local machine resources. I decided not even to test ByteDance’s UI-TARS for now, because it requested lower level access to my machine’s security and privacy features (if I test it out, I’ll definitely use a secondary computer). 

Testing reveals reasoning challenges

So the easiest to test are OpenAI’s Operator and Convergence’s Proxy. In our testing, the results highlighted how reasoning capabilities can matter more than raw automation features. Operator, in particular, was more buggy.

For example, I asked the agents to find and summarize VentureBeat’s five most popular stories. It was an ambiguous task, because VentureBeat doesn’t have a “most popular” section per se. Operator struggled with this. It first fell into an infinite scrolling loop while searching for ‘most popular’ stories, requiring manual intervention. In another attempt, it found a three-year-old article titled “Top five stories of the week.” In contrast, Proxy demonstrated better reasoning by identifying the five most visible stories on the homepage as a practical proxy for popularity, and it gave accurate summaries.

The distinction became even clearer in real-world tasks. I asked the agents to book a reservation at a romantic restaurant for noon in Napa, California. Operator approached the task linearly — finding a romantic restaurant first, then checking availability at noon. When no tables were available, it reached a dead end. Proxy showed more sophisticated reasoning by starting with OpenTable to find restaurants that were both romantic and available at the desired time. It even came back with a slightly better rated restaurant.

Even seemingly simple tasks revealed important differences. When searching for a “YubiKey 5C NFC price” on Amazon, Proxy quickly found the item more easily than Operator. 

OpenAI hasn’t divulged much about technologies it uses for training its Operator agent, other than saying it has trained its model on browser-use tasks. Convergence, however, has provided more detail: Its agent uses something called Generative Tree Search to “leverage Web-World Models that predict the state of the web after a proposed action has been taken. These are generated recursively to produce a tree of possible futures that are searched over to select the next optimal action, as ranked by our value models. Our Web-World models can also be used to train agents in hypothetical situations without generating a lot of expensive data.” (More here).

Benchmarks may be useless for now

On paper, these tools appear closely matched. Convergence’s Proxy achieves 88% on the WebVoyager benchmark, which evaluates web agents across 643 real-world tasks on 15 popular websites like Amazon and Booking.com. OpenAI’s Operator scores 87%, while Browser-Use says it reaches 89% but only after changing the WebVoyager codebase slightly, it conceded, “according to our needs”.

These benchmark scores should really be taken with a grain of salt, though, as they can be gamed. The real test comes in practical usage for real-world cases. It’s very early, the space is so rapidly changing, and these products are changing almost on a daily basis. The results will depend more on the specific jobs you’re trying to do, and you may want to instead rely on the vibes you get while using the different products.

Enterprise implications

The implications for enterprise automation are significant. As Witteveen points out in our video podcast conversation about this, where we do a deep dive into this browser-use trend, many companies are currently paying for virtual assistants – operated by real people – to handle basic web research and data gathering tasks. These browser-use agents could dramatically change that equation.

“If AI takes this over,” Witteveen notes, “that’s going to be some of the first low hanging fruit of people losing their jobs. It’s going to show up in some of these kinds of things.”

This could feed into the robotic process automation (RPA) trend, where browser use is pulled in as just another tool for companies to automate more tasks. And as mentioned earlier, the more powerful uses cases will be when an agent combined browser use with other tools, including things like Deep Research, where an LLM-driven agent uses a search tool plus browser use to do more sophisticated jobs.

Cost dynamics driving innovation

Another key factor driving rapid development is the availability of powerful open-source reasoning models like DeepSeek-R1. This allows companies building these browser-use agents to compete effectively with larger players by leveraging these models rather than building their own.

The pricing pressure is already evident. While OpenAI requires a $200 monthly ChatGPT Pro subscription to access Operator, Convergence offers limited free use (up to five uses per day) and a $20/month unlimited plan. This competitive dynamic should accelerate enterprise adoption, though clear use cases are still emerging.

Security and integration challenges

Several hurdles remain before widespread enterprise adoption. Some websites actively block automated browsing, while others require CAPTCHA verification. While OpenAI and Convergence have tools that can get past CAPTCHAs, they let users take over the task to fill them out — instead of doing them directly, since the whole point of CAPTCHAs is to ensure a human is at the other end. Tools like ByteDance’s UI-TARS request deep system access, which raises security concerns for enterprise deployment.

Additionally, the approach to website cooperation varies. OpenAI has worked with specific partners like Instacart, Priceline, DoorDash and Etsy, while others attempt to navigate any website. This inconsistency could impact reliability for enterprise use cases. And of course, any time an agent hits a site requiring login details, that will slow things — as the agents will turn things over to you to fill in those details.

Looking ahead

For enterprises evaluating these tools, the focus should be on specific use cases where autonomous web interaction could provide clear value – whether in research, customer service, or process automation. The technology is progressing rapidly, but success will depend on matching capabilities to concrete business needs.

As this space evolves, expect to see more enterprise-focused features and potentially specialized agents for specific industries or tasks. The race between established players and innovative startups should drive both technical advancement and competitive pricing, making 2025 a crucial year for enterprise browser-use agent adoption.

For more detail on these trends and testing results, check out the full video conversation between Sam Witteveen and myself.

Shape
Shape
Stay Ahead

Explore More Insights

Stay ahead with more perspectives on cutting-edge power, infrastructure, energy,  bitcoin and AI solutions. Explore these articles to uncover strategies and insights shaping the future of industries.

Shape

Energy Secretary Keeps Guam Power On

WASHINGTON—U.S. Secretary of Energy Chris Wright today issued an emergency order to mitigate the risk of blackouts for hardworking families and businesses in Guam. The emergency order permits the Guam Power Authority (GPA) to operate specified generation units to meet anticipated electricity demand and maintain reliability. The order follows a request from GPA. “President Trump and the Department of Energy remain committed to using every available tool to reduce the risk of power outages and lower energy costs for hardworking families and businesses,” said Secretary Wright. “Today’s order responds to the urgent need to strengthen grid reliability while ensuring the people of Guam have access to affordable, reliable, and secure electricity.” GPA has limited generation options to meet demand due to limitations, including but not limited to, constraints with the Ukudu steam turbine going offline for emergency repairs. This order is effective 10:00 AM Guam Standard Time on August 6, 2026, through 11:59 PM Guam Standard Time on November 4, 2026.                                                                                            ###

Read More »

DOE’s Office of Energy Dominance Financing Closes Loan to Strengthen Puerto Rico’s Grid, Delivering Hundreds of Millions in Electricity Cost Savings

WASHINGTON – The U.S. Department of Energy’s (DOE) Office of Energy Dominance Financing (EDF) today announced it has closed a $489.4 million loan to Amanecer Puerto Rico LLC, a subsidiary of Pattern Energy, to lower electricity costs and strengthen Puerto Rico’s electric grid. Thanks to President Trump’s Working Families Tax Cuts Act, the investment is expected to save Puerto Rican families and businesses approximately $312.5 million in electricity costs over the next 25 years while improving grid reliability, strengthening energy security, supporting American manufacturing, and reducing dependence on foreign-controlled supply chains. “President Trump’s Working Families Tax Cuts Act is driving investments that strengthen America’s energy infrastructure while lowering costs for hardworking families,” said EDF Director Gregory A. Beard. “This investment will strengthen Puerto Rico’s electric grid, lower electricity costs, support American manufacturing, and provide a pathway for the reliable, dispatchable power needed to deliver affordable, reliable, and secure energy for the people of Puerto Rico.” Puerto Rico’s grid has experienced chronic outages and prolonged service interruptions that have imposed significant costs on families, businesses, and critical infrastructure. Following a comprehensive review by the Trump Administration, DOE restructured the financing to better align with the Administration’s priorities of lowering energy costs, strengthening American manufacturing, and ensuring the deployment of reliable, secure energy infrastructure. The financing will support: 220 megawatts of battery energy storage systems in Arecibo and Santa Isabel using American-manufactured battery technology and secure domestic supply chains. Battery storage capable of providing backup electricity for more than 100,000 customers during power shortages and helping avoid approximately 13 million customer interruption hours based on 2025 operating data. A pathway for the future development of reliable, dispatchable natural gas-fired generation to improve grid stability and strengthen Puerto Rico’s long-term energy security. This financing builds on the Trump Administration’s broader efforts to restore

Read More »

United States to Host International Atomic Energy Agency Launch for New Maritime Nuclear Initiative

WASHINGTON—The United States will host the ministerial launch of the International Atomic Energy Agency’s (IAEA) new initiative, Atomic Technologies Licensed for Applications at Sea (ATLAS), in Washington, D.C., on August 26–27, 2026. This landmark event will bring together ministers, policymakers, and industry leaders from around the world to advance the safe and secure use of nuclear technologies in the maritime sector. The ATLAS initiative aims to create an international framework to address legal and regulatory complexities to enable the deployment of nuclear applications at sea. It builds upon the IAEA’s extensive experience and global authority in nuclear safety, security, and safeguards to support the deployment of innovative civil nuclear technologies. The initiative is technology-neutral and focuses on establishing and maintaining the highest standards for safety, security, and nonproliferation. “DOE remains focused on unleashing American energy dominance, accelerating innovation, and advancing sources of energy that are affordable, reliable, and secure for the American people,” said U.S. Secretary of Energy Chris Wright. “Hosting the launch of ATLAS supports this critical mission, positioning the U.S. nuclear and maritime sectors at the forefront of advanced energy innovation, while promoting safety and security for the United States and the world.”  Leading up to the launch, the United States will hold an “Industry Day” on August 25, 2026. This event will serve as a platform for America’s leading nuclear and maritime companies to showcase cutting-edge technologies. The Industry Day will highlight American innovation and underscore American energy dominance that will support the global expansion of civil maritime nuclear applications. “The global maritime sector is at a critical turning point, facing urgent pressure to sustain long-distance, high-speed operations while ensuring reliability and energy security,” said IAEA Director General Rafael Mariano Grossi. “Nuclear energy is fast emerging as a game-changer. Small modular reactors offer a safe and viable option for

Read More »

Vista seeks RIGI approval for $5.8 billion Bandurria Norte development

Vista Energy SAB de CV has applied to include its Bandurria Norte shale oil development in Argentina’s Large Investment Incentive Regime (RIGI). The $5.8-billion project targets peak production of 50,000 boe/d. Bandurria Norte is the largest oil project submitted under RIGI, which provides tax, customs, and foreign-exchange incentives for qualifying investments, and follows approval of Pampa Energía SA’s $4.521 billion Rincón de Aranda development, which established the first framework for qualifying incremental shale oil production under the regime. Together, the projects represent more than $10.3 billion in planned investment and would extend RIGI-backed development into undeveloped Vaca Muerta oil acreage. Bandurria Norte spans 26,500 acres in Vaca Muerta’s oil window and currently has no producing wells. Vista plans to drill and complete 332 horizontal wells and build dedicated infrastructure, including a 40,000-b/d oil treatment plant, a gas compression plant, gathering systems, pipelines, and associated infrastructure. The project would be Vista’s first large-scale development outside its core Bajada del Palo hub, where existing roads, processing capacity, and gathering networks support lower-cost drilling. Bandurria Norte requires full greenfield development, increasing upfront capital requirements, and execution risk. Vista said RIGI incentives are material to developing its undeveloped acreage because incremental production from new areas can qualify separately from existing output if volumes remain physically and operationally traceable. Export capacity remains critical Bandurria Norte forms part of Vista’s plan to increase production to 208,000 boe/d in 2028 and 250,000 boe/d in 2030. Development depends on additional crude transportation capacity from the Neuquén basin, particularly the Vaca Muerta Oil Sur (VMOS) pipeline under construction between Allen and Punta Colorada in Río Negro province. Designed for an initial capacity of 550,000 b/d and expandable to 700,000 b/d, VMOS is scheduled for start-up in first-half 2027. Vista averaged 156,000 boe/d of production in second-quarter 2026, up 16%

Read More »

Oil prices surge on renewed Middle East tensions

@import url(‘https://fonts.googleapis.com/css2?family=Inter:wght@100..900&display=swap’); .ebm-page__main h1, .ebm-page__main h2, .ebm-page__main h3, .ebm-page__main h4, .ebm-page__main h5, .ebm-page__main h6 { font-family: Inter; } body { line-height: 150%; letter-spacing: 0.025em; } button, .ebm-button-wrapper { font-family: Inter; } .label-style { text-transform: uppercase; color: var(–color-grey); font-weight: 600; font-size: 0.75rem; } .caption-style { font-size: 0.75rem; opacity: .6; } #onetrust-pc-sdk [id*=btn-handler], #onetrust-pc-sdk [class*=btn-handler] { background-color: #c19a06 !important; border-color: #c19a06 !important; } #onetrust-policy a, #onetrust-pc-sdk a, #ot-pc-content a { color: #c19a06 !important; } #onetrust-consent-sdk #onetrust-pc-sdk .ot-active-menu { border-color: #c19a06 !important; } #onetrust-consent-sdk #onetrust-accept-btn-handler, #onetrust-banner-sdk #onetrust-reject-all-handler, #onetrust-consent-sdk #onetrust-pc-btn-handler.cookie-setting-link { background-color: #c19a06 !important; border-color: #c19a06 !important; } #onetrust-consent-sdk .onetrust-pc-btn-handler { color: #c19a06 !important; border-color: #c19a06 !important; } Global oil prices rallied sharply July 29 as renewed military escalation in the Middle East ended several days of relative calm and revived concerns over crude flows from the region. Brent crude surged 7% to above $90/bbl, while US WTI climbed above $84/bbl. The rally followed joint US-Saudi airstrikes on Iran-backed militias in Iraq—which killed at least 20 fighters, according to Iraq’s Popular Mobilization Forces—and a retaliatory Iranian missile barrage targeting US forces in the region. Stay updated on oil price volatility, shipping disruptions, LNG market analysis, and production output through OGJ’s Iran war content hub. This operation marks the first time Saudi Arabia has publicly acknowledged a combat role in the conflict. Washington and Riyadh stated that the strikes were launched in response to drone attacks on oil facilities in Saudi Arabia’s Eastern Province—attacks that originated from within Iraq. The sudden escalation across multiple fronts has raised concerns that the 5-month-old conflict could expand further, threatening critical shipping lanes—the Strait of Hormuz and, following the Houthis’ declared blockade of Saudi shipping, the Bab el-Mandeb strait. Adding support to prices, US commercial crude inventories fell by 7.2 million bbl in the week ended July 24, according

Read More »

Türkiye signs partnership deal with bp for Kirkuk oil field redevelopment

bp plc has farmed out a 15% interest in BP Energy Co. of Kirkuk Ltd. (BP ECKL) to state-owned Turkish Petroleum Corp. (TPAO), expanding on a partnership to support the redevelopment of major oil and gas fields in the Kirkuk region of northern Iraq. The move comes as Iraq aims to increase oil and gas production through various international partnerships. Signed during the official visit of Iraqi Prime Minister Ali Al-Zaidi to Türkiye, the agreement builds on a strategic cooperation memorandum of understanding (MoU) signed by the companies in February 2026.  The development and production contract covers an initial phase of oil and gas production of more than 3 billion boe from the Baba and Avanah domes of Kirkuk oil field and the adjacent Bai Hassan, Jambur, and Khabbaz fields in Federal Iraq, all currently operated by the North Oil Co. (NOC) and North Gas Co. (NGC), bp said in a release July 28. The contract area holds potential for additional exploration, the companies said. The deal follows one that saw ConocoPhillips acquire a 42% interest in BP ECKL. Together, bp said, the transactions support the next phase of redevelopment in Kirkuk. Türkiye Energy and Natural Resources Minister Bayraktar said the agreement is a step “towards making Turkish Petroleum a company that produces 1 million barrels of oil and natural gas per day.” Following completion of the transaction, which is subject to regulatory approvals, bp will remain the majority shareholder and a key participant in BP ECKL (bp 43%, ConocoPhillips 42%, TPAO 15%).

Read More »

Data center energy constraints and moratoriums are mounting. Expect to see stalled AI projects

Fuel cells are more efficient, he says, and don’t emit particulates, but they’re more expensive and less reliable than generators and have other operational issues. One company that recently decided to go with fuel cells is Oracle, which will use 2.45 gigawatts worth of fuel cells to power its Project Jupiter data center in New Mexico, replacing the previous plan to use gas turbines and diesel generators. According to Oracle, the fuel cells will significantly reduce emissions, use only a “negligible” amount of water, and be quieter than turbines and generators. Plus, the on-site power generation will help protect energy rates of area residents. “If you want to build a data center, there’s a better way to build it,” says Natalie Sunderland, chief marketing and communications officer at Bloom Energy, which makes the fuel cells that Oracle plans to deploy. And companies aren’t about to scale back on their AI ambitions or reduce their demands for data centers, she says.

Read More »

Nvidia’s Next Move? Financing AI

The report argues that AI infrastructure spending is on track to exceed $2 trillion annually by 2028, with cumulative investment reaching roughly $11.1 trillion between 2024 and 2029. Financing those projects will require a massive expansion of credit markets, which could result in a cumulative collective AI-related debt of $7 trillion by the end of the decade. With deals reaching into the multibillion-dollar range, banks and venture funds simply don’t have that kind of money. Enter Nvidia. As of the first fiscal quarter of 2027 ended April 26, 2026, Nvidia was sitting on roughly $80.5 billion in cash, cash equivalents, and short-term investments. After data center capacity constrained AI expansion in 2025 and chip supply became the limiting factor in early 2026, financing has emerged as the next major obstacle to scaling AI infrastructure. “It is clear that financing will now be one of the most significant obstacles to ramping large-scale compute broadly available to everyone,” the authors wrote.

Read More »

The Gigawatt Buildout Faces the Execution Test

Demand remains abundant across the data center and AI infrastructure market. Alphabet has raised its capital spending forecast again. OpenAI has unveiled a 3.2-gigawatt project in Georgia. Hut 8 has signed another multibillion-dollar lease in Texas, while BlackRock and its partners have acquired Aligned Data Centers for approximately $40 billion. Oracle, meanwhile, could face a $7 billion collateral requirement in Wisconsin, Meta is paying more to finance a $12 billion Texas project, and proposed developments are drawing resistance across North America and beyond. Together, these developments describe a market moving into a more demanding phase. Land, power and capital remain available, but not on the same terms everywhere. The projects most likely to move are those that combine a credible customer, a durable power path, institutional financing and a development strategy capable of surviving public scrutiny. Hyperscaler Spending Keeps Rising After Google Cloud revenue grew 82% year over year to $24.8 billion in the second quarter, Alphabet increased its expected 2026 capital expenditures to between $195 billion and $205 billion. That was $15 billion above its previous forecast and more than twice the $91.5 billion the company spent in 2025. The cloud growth explains the spending without removing its financial pressure. The largest platforms are committing unprecedented cash to facilities, chips, networks and power systems whose economics will be measured over many years. OpenAI’s Project Camellia in Effingham County, Georgia, extends the scale further. OpenAI says it is designing and developing the campus itself and has contracted with Georgia Power for 3.2 GW to be delivered in phases from 2028 through 2032. Local reporting places the initial investment at at least $20 billion across roughly 1,400 acres. The 25-year power supply agreement may be as important as the campus size. OpenAI says it will cover the infrastructure costs, protect residential

Read More »

AI Clusters and the New Economics of Data Center Optics: A Conversation with Cisco’s Bill Gartner

Three Networks Inside the AI Factory Understanding the optical challenge begins with recognizing that an AI cluster contains several distinct networking environments. Gartner divided AI infrastructure into three broad tiers: scale-up, scale-out and scale-across. Each operates over a different distance, carries a different level of traffic and creates a different set of requirements for the interconnect. Scale-up describes the connections within a rack, where operators place as many GPUs as possible inside servers and then pack those servers into the available rack footprint. Gartner estimated that the bandwidth within this environment can be approximately 500 times that of a traditional wide-area network application. Scale-up connections still rely heavily on electrical interfaces because electrical interconnects remain relatively inexpensive and power efficient over short distances. Once the compute capacity of a rack has been exhausted, the cluster must expand into additional racks. This is the scale-out network, where 400G and 800G pluggable optics connect large numbers of GPU systems operating in parallel. Gartner characterized scale-out bandwidth as roughly 50 times the capacity associated with a conventional WAN environment. The third tier, scale-across, emerges when a data center reaches its practical power limit and the AI infrastructure must extend into another facility. Those data centers may be separated by tens or hundreds of kilometers, requiring coherent optical technology capable of carrying extremely high-capacity signals over longer distances. Scale-across networks can represent approximately 14 times traditional WAN bandwidth, according to Gartner. Taken together, the three tiers illustrate why optics has become inseparable from the AI infrastructure discussion. Network capacity must expand inside the rack, across rows of racks and increasingly between separate data centers—all without consuming an untenable share of the power budget or introducing failures that leave GPUs idle. When One Link Slows the Whole Cluster The reliability requirement for AI networks differs

Read More »

Meta’s Canadian AI Data Center: A New Model for Infrastructure and Energy Integration

Canada’s expansion as an artificial intelligence infrastructure market received its strongest endorsement yet on July 8, when Meta broke ground on a data center campus representing an investment of more than C$13 billion. The project, located in Sturgeon County north of Edmonton, will be Meta’s first data center in Canada and the 33rd facility in its global portfolio. Planned initially at 1 GW of power capacity, the AI-optimized campus could eventually scale to 1.8 GW, placing it among the largest data center developments under construction anywhere outside the United States. The announcement follows the Canadian federal government’s May launch of consultations on a forthcoming National Electricity Strategy, which identifies AI data centers as a major source of future electricity demand.  Approximately 3,000 construction workers are expected to be on the site at peak activity, while more than 300 permanent employees will operate the campus after completion. Meta is also committing approximately C$60 million to improvements involving local roads, water systems and other community infrastructure. The Meta announcement illustrates a fundamental change in Canadian data center construction. Rather than selecting a building site and applying for an ordinary utility connection, Meta and its partners have spent years coordinating the data center with a purpose-built 932 MW generating station, grid upgrades and long-term natural-gas transportation agreements. The project effectively combines a data center, a power plant and an infrastructure development program into a single construction ecosystem. Construction Has Begun on the Canadian Campus Meta describes the Sturgeon County project as an AI-optimized facility intended to support the computing demands of its core platforms, AI services and connected devices. The company said the buildout will occur in phases rather than delivering the entire gigawatt at once. Alberta’s major-project registry estimates a roughly three-year construction period. That timeline will require a sustained deployment of

Read More »

Nuclear Momentum Meets the Megawatt Test

Valar then supplied the most visible connection between the criticality program and data center technology. After reaching criticality, Valar advanced Ward 250 to approximately 10 kilowatts of thermal output and conducted a separate demonstration in which power from the reactor was used to run Nvidia Blackwell-based computing hardware. On July 1, Valar and Nvidia also announced that they were exploring a small Utah data center using closed-loop cooling and behind-the-meter advanced nuclear generation. The demonstration load was microscopic beside a hyperscale campus that may require hundreds of megawatts. Nvidia described the work as an exploration of how behind-the-meter advanced nuclear systems could support future AI factories, not as an agreement to purchase a specified quantity of electricity. Deployable Energy became the third developer to achieve zero-power criticality when its Unity reactor completed its experiment at Idaho National Laboratory on June 30. DOE announced the result July 1, noting that the three companies had satisfied the administration’s objective of achieving three advanced reactor criticality milestones by July 4. The commercial follow-up came quickly. On July 7, Deployable Energy and energy-infrastructure facilitator GridMarket announced a partnership aimed at data centers, hyperscalers and industrial customers. The agreement includes a committed pilot project and priority access to future Unity capacity. The companies said they were targeting 500 megawatts of annual deployments from 2030 through 2035 and more than 3 gigawatts cumulatively. The companies have not publicly named the pilot host or end customers. Even so, the committed pilot and access provisions put the arrangement ahead of a conventional memorandum of understanding. GridMarket is attempting to assemble sites, customers, technology and capital before commercial Unity units become available. Aalo Atomics completed the fourth criticality experiment on July 4, with DOE announcing the achievement July 6. Aalo-X went from groundbreaking to a sustained chain reaction in

Read More »

Microsoft will invest $80B in AI data centers in fiscal 2025

And Microsoft isn’t the only one that is ramping up its investments into AI-enabled data centers. Rival cloud service providers are all investing in either upgrading or opening new data centers to capture a larger chunk of business from developers and users of large language models (LLMs).  In a report published in October 2024, Bloomberg Intelligence estimated that demand for generative AI would push Microsoft, AWS, Google, Oracle, Meta, and Apple would between them devote $200 billion to capex in 2025, up from $110 billion in 2023. Microsoft is one of the biggest spenders, followed closely by Google and AWS, Bloomberg Intelligence said. Its estimate of Microsoft’s capital spending on AI, at $62.4 billion for calendar 2025, is lower than Smith’s claim that the company will invest $80 billion in the fiscal year to June 30, 2025. Both figures, though, are way higher than Microsoft’s 2020 capital expenditure of “just” $17.6 billion. The majority of the increased spending is tied to cloud services and the expansion of AI infrastructure needed to provide compute capacity for OpenAI workloads. Separately, last October Amazon CEO Andy Jassy said his company planned total capex spend of $75 billion in 2024 and even more in 2025, with much of it going to AWS, its cloud computing division.

Read More »

John Deere unveils more autonomous farm machines to address skill labor shortage

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Self-driving tractors might be the path to self-driving cars. John Deere has revealed a new line of autonomous machines and tech across agriculture, construction and commercial landscaping. The Moline, Illinois-based John Deere has been in business for 187 years, yet it’s been a regular as a non-tech company showing off technology at the big tech trade show in Las Vegas and is back at CES 2025 with more autonomous tractors and other vehicles. This is not something we usually cover, but John Deere has a lot of data that is interesting in the big picture of tech. The message from the company is that there aren’t enough skilled farm laborers to do the work that its customers need. It’s been a challenge for most of the last two decades, said Jahmy Hindman, CTO at John Deere, in a briefing. Much of the tech will come this fall and after that. He noted that the average farmer in the U.S. is over 58 and works 12 to 18 hours a day to grow food for us. And he said the American Farm Bureau Federation estimates there are roughly 2.4 million farm jobs that need to be filled annually; and the agricultural work force continues to shrink. (This is my hint to the anti-immigration crowd). John Deere’s autonomous 9RX Tractor. Farmers can oversee it using an app. While each of these industries experiences their own set of challenges, a commonality across all is skilled labor availability. In construction, about 80% percent of contractors struggle to find skilled labor. And in commercial landscaping, 86% of landscaping business owners can’t find labor to fill open positions, he said. “They have to figure out how to do

Read More »

2025 playbook for enterprise AI success, from agents to evals

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More 2025 is poised to be a pivotal year for enterprise AI. The past year has seen rapid innovation, and this year will see the same. This has made it more critical than ever to revisit your AI strategy to stay competitive and create value for your customers. From scaling AI agents to optimizing costs, here are the five critical areas enterprises should prioritize for their AI strategy this year. 1. Agents: the next generation of automation AI agents are no longer theoretical. In 2025, they’re indispensable tools for enterprises looking to streamline operations and enhance customer interactions. Unlike traditional software, agents powered by large language models (LLMs) can make nuanced decisions, navigate complex multi-step tasks, and integrate seamlessly with tools and APIs. At the start of 2024, agents were not ready for prime time, making frustrating mistakes like hallucinating URLs. They started getting better as frontier large language models themselves improved. “Let me put it this way,” said Sam Witteveen, cofounder of Red Dragon, a company that develops agents for companies, and that recently reviewed the 48 agents it built last year. “Interestingly, the ones that we built at the start of the year, a lot of those worked way better at the end of the year just because the models got better.” Witteveen shared this in the video podcast we filmed to discuss these five big trends in detail. Models are getting better and hallucinating less, and they’re also being trained to do agentic tasks. Another feature that the model providers are researching is a way to use the LLM as a judge, and as models get cheaper (something we’ll cover below), companies can use three or more models to

Read More »

OpenAI’s red teaming innovations define new essentials for security leaders in the AI era

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More OpenAI has taken a more aggressive approach to red teaming than its AI competitors, demonstrating its security teams’ advanced capabilities in two areas: multi-step reinforcement and external red teaming. OpenAI recently released two papers that set a new competitive standard for improving the quality, reliability and safety of AI models in these two techniques and more. The first paper, “OpenAI’s Approach to External Red Teaming for AI Models and Systems,” reports that specialized teams outside the company have proven effective in uncovering vulnerabilities that might otherwise have made it into a released model because in-house testing techniques may have missed them. In the second paper, “Diverse and Effective Red Teaming with Auto-Generated Rewards and Multi-Step Reinforcement Learning,” OpenAI introduces an automated framework that relies on iterative reinforcement learning to generate a broad spectrum of novel, wide-ranging attacks. Going all-in on red teaming pays practical, competitive dividends It’s encouraging to see competitive intensity in red teaming growing among AI companies. When Anthropic released its AI red team guidelines in June of last year, it joined AI providers including Google, Microsoft, Nvidia, OpenAI, and even the U.S.’s National Institute of Standards and Technology (NIST), which all had released red teaming frameworks. Investing heavily in red teaming yields tangible benefits for security leaders in any organization. OpenAI’s paper on external red teaming provides a detailed analysis of how the company strives to create specialized external teams that include cybersecurity and subject matter experts. The goal is to see if knowledgeable external teams can defeat models’ security perimeters and find gaps in their security, biases and controls that prompt-based testing couldn’t find. What makes OpenAI’s recent papers noteworthy is how well they define using human-in-the-middle

Read More »