Stay Ahead, Stay ONMINE

The rise of browser-use agents: Why Convergence’s Proxy is beating OpenAI’s Operator

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More A new wave of AI-powered browser-use agents is emerging, promising to transform how enterprises interact with the web. These agents can autonomously navigate websites, retrieve information, and even complete transactions – but early testing reveals significant […]

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More


A new wave of AI-powered browser-use agents is emerging, promising to transform how enterprises interact with the web. These agents can autonomously navigate websites, retrieve information, and even complete transactions – but early testing reveals significant gaps between promise and performance.

While consumer examples offered by OpenAI’s new browser-use agent Operator, like ordering pizza or buying game tickets, have grabbed headlines, the question is about where the main developer and enterprise use cases are. “The thing that we don’t know is what will be the killer app,” said Sam Witteveen, co-founder of Red Dragon, a company that develops AI agent applications. “My guess is it’s going to be things that just take time on the web that you don’t actually enjoy.” This includes things like going on the web and searching for the cheapest price of a product or booking the best hotel accommodations. More likely it will be used in combination with other tools like Deep Research, where companies can then do even more sophisticated research plus execution of tasks around the web.

Companies need to carefully evaluate the rapidly evolving landscape as established players and startups take different approaches to solving the autonomous browsing challenge.

Key players in the browser-use agent landscape

The field has quickly become crowded with both major tech companies and innovative startups:

Operator and Proxy are the most advanced, in terms of being consumer-friendly and out-of-the-box ready. Many of the others appear to be positioning themselves more for developer or enterprise usage. For example, Browser Use, a Y-Combinator startup that allows users to customize the models used with the agent. This gives you more control over how the agent works, including using a model from your local machine. But it’s definitely more involved.

The others listed above provide a varying degree of functionality and interaction with local machine resources. I decided not even to test ByteDance’s UI-TARS for now, because it requested lower level access to my machine’s security and privacy features (if I test it out, I’ll definitely use a secondary computer). 

Testing reveals reasoning challenges

So the easiest to test are OpenAI’s Operator and Convergence’s Proxy. In our testing, the results highlighted how reasoning capabilities can matter more than raw automation features. Operator, in particular, was more buggy.

For example, I asked the agents to find and summarize VentureBeat’s five most popular stories. It was an ambiguous task, because VentureBeat doesn’t have a “most popular” section per se. Operator struggled with this. It first fell into an infinite scrolling loop while searching for ‘most popular’ stories, requiring manual intervention. In another attempt, it found a three-year-old article titled “Top five stories of the week.” In contrast, Proxy demonstrated better reasoning by identifying the five most visible stories on the homepage as a practical proxy for popularity, and it gave accurate summaries.

The distinction became even clearer in real-world tasks. I asked the agents to book a reservation at a romantic restaurant for noon in Napa, California. Operator approached the task linearly — finding a romantic restaurant first, then checking availability at noon. When no tables were available, it reached a dead end. Proxy showed more sophisticated reasoning by starting with OpenTable to find restaurants that were both romantic and available at the desired time. It even came back with a slightly better rated restaurant.

Even seemingly simple tasks revealed important differences. When searching for a “YubiKey 5C NFC price” on Amazon, Proxy quickly found the item more easily than Operator. 

OpenAI hasn’t divulged much about technologies it uses for training its Operator agent, other than saying it has trained its model on browser-use tasks. Convergence, however, has provided more detail: Its agent uses something called Generative Tree Search to “leverage Web-World Models that predict the state of the web after a proposed action has been taken. These are generated recursively to produce a tree of possible futures that are searched over to select the next optimal action, as ranked by our value models. Our Web-World models can also be used to train agents in hypothetical situations without generating a lot of expensive data.” (More here).

Benchmarks may be useless for now

On paper, these tools appear closely matched. Convergence’s Proxy achieves 88% on the WebVoyager benchmark, which evaluates web agents across 643 real-world tasks on 15 popular websites like Amazon and Booking.com. OpenAI’s Operator scores 87%, while Browser-Use says it reaches 89% but only after changing the WebVoyager codebase slightly, it conceded, “according to our needs”.

These benchmark scores should really be taken with a grain of salt, though, as they can be gamed. The real test comes in practical usage for real-world cases. It’s very early, the space is so rapidly changing, and these products are changing almost on a daily basis. The results will depend more on the specific jobs you’re trying to do, and you may want to instead rely on the vibes you get while using the different products.

Enterprise implications

The implications for enterprise automation are significant. As Witteveen points out in our video podcast conversation about this, where we do a deep dive into this browser-use trend, many companies are currently paying for virtual assistants – operated by real people – to handle basic web research and data gathering tasks. These browser-use agents could dramatically change that equation.

“If AI takes this over,” Witteveen notes, “that’s going to be some of the first low hanging fruit of people losing their jobs. It’s going to show up in some of these kinds of things.”

This could feed into the robotic process automation (RPA) trend, where browser use is pulled in as just another tool for companies to automate more tasks. And as mentioned earlier, the more powerful uses cases will be when an agent combined browser use with other tools, including things like Deep Research, where an LLM-driven agent uses a search tool plus browser use to do more sophisticated jobs.

Cost dynamics driving innovation

Another key factor driving rapid development is the availability of powerful open-source reasoning models like DeepSeek-R1. This allows companies building these browser-use agents to compete effectively with larger players by leveraging these models rather than building their own.

The pricing pressure is already evident. While OpenAI requires a $200 monthly ChatGPT Pro subscription to access Operator, Convergence offers limited free use (up to five uses per day) and a $20/month unlimited plan. This competitive dynamic should accelerate enterprise adoption, though clear use cases are still emerging.

Security and integration challenges

Several hurdles remain before widespread enterprise adoption. Some websites actively block automated browsing, while others require CAPTCHA verification. While OpenAI and Convergence have tools that can get past CAPTCHAs, they let users take over the task to fill them out — instead of doing them directly, since the whole point of CAPTCHAs is to ensure a human is at the other end. Tools like ByteDance’s UI-TARS request deep system access, which raises security concerns for enterprise deployment.

Additionally, the approach to website cooperation varies. OpenAI has worked with specific partners like Instacart, Priceline, DoorDash and Etsy, while others attempt to navigate any website. This inconsistency could impact reliability for enterprise use cases. And of course, any time an agent hits a site requiring login details, that will slow things — as the agents will turn things over to you to fill in those details.

Looking ahead

For enterprises evaluating these tools, the focus should be on specific use cases where autonomous web interaction could provide clear value – whether in research, customer service, or process automation. The technology is progressing rapidly, but success will depend on matching capabilities to concrete business needs.

As this space evolves, expect to see more enterprise-focused features and potentially specialized agents for specific industries or tasks. The race between established players and innovative startups should drive both technical advancement and competitive pricing, making 2025 a crucial year for enterprise browser-use agent adoption.

For more detail on these trends and testing results, check out the full video conversation between Sam Witteveen and myself.

Shape
Shape
Stay Ahead

Explore More Insights

Stay ahead with more perspectives on cutting-edge power, infrastructure, energy,  bitcoin and AI solutions. Explore these articles to uncover strategies and insights shaping the future of industries.

Shape

California Resources unloads Uinta assets

The leaders of California Resources Corp., Long Beach, have sold the company’s Uinta basin assets for about $90 million to an undisclosed buyer. The deal has an effective date of July 1 and is expected to close by yearend. “Today’s transaction strengthens our business,” said Francisco Leon, CRC president and chief executive officer. “This transaction enhances our capital allocation flexibility, allowing us to invest in higher-return opportunities within the Golden State, and supports our shareholder return strategy.” CRC had come to own the Uinta assets, which span about 100,000 net acres, after it acquired Berry Corp. in December of last year for $709 million. But the operation accounts for a small part of CRC’s business–2.5% of oil production and 8% of natural gas production in the second quarter–and Leon last month told analysts “it’s hard to see allocating a lot of dollars back into the Uinta” as his team focuses on building out its California network of assets. “It requires a pretty significant amount of capital to develop the scale that we need for a second asset,” Leon said Aug. 10 after CRC reported its second-quarter results. “So as we do a side-by-side and we compare the Uinta assets with California, Uinta has higher capital intensity, higher break-evens, lower crude quality [and] higher transportation and operating costs and steeper declines.” In the deal announcement, Leon said the Uinta sale also offsets the price CRC will pay for a set of midstream assets in California it plans to buy from CorEnergy Infrastructure Trust. The purchase of those pipelines and other operations is expected to close later this month. Shares of CRC (Ticker: CRC) were down slightly to $54.24 in late-morning trading Sept. 17. They have lost about 15% of their value over the past 6 months, trimming the company’s market capitalization

Read More »

Vitesse acquires interest in Chevron-operated DJ basin assets

Vitesse Energy Inc. has acquired non-operated oil and gas assets in the Denver-Julesburg (DJ) Basin in Colorado from an undisclosed seller for an initial unadjusted purchase price of $26 million. The acquired assets (average working interest: 4.1%), which lie primarily in Weld County, Colorado, are entirely operated by Chevron, and add “a high-quality, predominantly proved developed producing asset base,” said Jamie Benard, Vitesse’s chief executive officer and president, in a release Sept. 16. Vitesse said the deal adds to the company’s existing DJ basin position under a top-tier operator and adds 819 gross wells to its database, strengthening underwriting of future opportunities in the basin. Over the next 12 months following the effective date (June 1, 2026), the acquired assets are expected to produce about 900 boe/d on a two-stream basis (28% oil), the company said. In connection with the acquisition, the company has entered into commodity derivative contracts covering a significant portion of the acquired production through 2030 to support the underwritten returns. Prior to the deal closing Sept. 15, 2026, the company noted in an August 2026 investor presentation that it holds fractional, non-operated working interests in productive wells and new drills across the Williston, Powder River, and DJ basins with an average 3.6% average working interest.

Read More »

Continental Resources, PDVSA sign MoU for potential Venezuela oil development

Continental Resources Inc., Oklahoma City, Okla., has signed a Memorandum of Understanding (MOU) with Petróleos de Venezuela SA (PDVSA) to operate and develop the Ayacucho 2 Block in Venezuela’s Orinoco Oil Belt. In a release Sept. 15, Continental said it has the opportunity “to bring significant private capital, technology, technical expertise, and large-scale operating capability to the redevelopment of Venezuela’s oil industry.” The Ayacucho 2 Block lies north of the Orinoco River in Anzoátegui state. The 126,000-acre block contains an estimated 30 billion bbl of resource in place, Continental said. Upon signing a long-term Contrato de Participación Productiva (CPP) agreement—expected in the coming weeks—Continental would operate the block with a 100% working interest, it said in a release Sept. 16. Continental Resources has been building its international presence in recent years, including in Türkiye’s Diyarbakır Basin and Argentina’s Vaca Muerta formation.

Read More »

EIA: US crude inventories down 600,000 bbl

US crude oil inventories for the week ended Sept. 11, excluding the Strategic Petroleum Reserve, decreased by 600,000 bbl from the previous week, according to data from the US Energy Information Administration (EIA). At 423.4 million bbl, US crude oil inventories are 1% above the 5-year average for this time of year, the EIA report indicated. EIA said total motor gasoline inventories increased by 800,000 bbl from last week and are 5% below the 5-year average for this time of year. Distillate inventories increased by 1.6 million bbl last week and are about 13% below the 5-year average for this time of year. Propane-propylene inventories decreased 1.4 million bbl, 22% above the 5-year average. Total commercial petroleum inventories increased by 2.6 million bbl for the week. US refineries processed 17.3 million b/d for the week ended Sept. 11, which was 256,000 b/d less than the previous week’s average. Refineries operated at 96.8% of capacity. Gasoline output averaged 9.6 million b/d, and distillate production decreased to 5.2 million b/d. US crude oil imports averaged 7.1 million b/d, up 234,000 b/d from the previous week. Over the last 4 weeks, crude oil imports averaged about 6.7 million b/d, 7.5% more than the same 4-week period last year. Total motor gasoline imports averaged 537,000 b/d. Distillate fuel imports averaged 114,000 b/d.

Read More »

POSCO to acquire Chord’s Marcellus gas assets for $550 million

South Korea-based POSCO International Corp. has agreed to acquire the non-operated Marcellus position of a Chord Energy Corp. subsidiary for $550 million, gaining a producing US shale gas asset that it plans to use to generate immediate cash flow while expanding its LNG value chain. The acquisition includes about 32,000 net acres in the core of the Marcellus play in Pennsylvania, trailing 12-month production of about 121 MMcfd, and 1.3 tcf of reserves, including 220 bcf of discovered potential, POSCO said in a briefing Sept. 16. The asset produces 100% residue gas with no NGLs and includes 2,006 wells, consisting of 1,305 producing wells and 701 development wells. POSCO said first-half 2026 production averaged about 124 MMscfd. Field activities, including production, drilling, and permitting, will continue to be managed by an established local operator, POSCO said. The company plans to focus on gas marketing and downstream integration. “This investment goes beyond the simple acquisition of a producing gas field,” said Dong-il Kim, head of POSCO International’s E&P Business Division. “It is an investment to expand the value chain by securing immediate returns through proven US upstream assets and connecting gas sales, liquefaction, LNG trading, and group demand.” POSCO said that, under the current sales portfolio, about half of production is sold near production sites, with roughly 30% marketed into northeastern US and Ohio and another 20% supplied to Gulf Coast markets. Beginning in 2029, the company plans to direct a portion of production to LNG liquefaction plants and market the resulting LNG through its trading subsidiary. 

Read More »

Plains to acquire Powder River Basin assets from Silver Creek for $585 million

Plains All American Pipeline LP and Plains GP Holdings, through a subsidiary, have agreed to acquire SCM PR II LLC (Silver Creek) from subsidiaries of Tailwater Capital and The Energy and Minerals Group for about $585 million in cash. The transaction will expand Plains’ Powder River Basin footprint, increase connectivity to producer supply and strengthen its Rockies crude oil gathering and transportation network, the company said in a release Sept. 16. “This is a highly strategic addition to Plains’ Rockies platform,” said Willie Chiang, chairman, chief executive officer, and president of Plains. “It expands our footprint in the Powder River Basin, a region supported by substantial remaining drilling inventory and an attractive outlook for continued producer development over the coming years.” Silver Creek owns and operates a large integrated crude oil gathering system in the Powder River Basin across Converse, Campbell, Johnson, and Natrona counties, Wyoming, providing producers access to Plains’ existing Rockies infrastructure through the Guernsey and Fort Laramie hubs. The acquired assets include about 600 miles of crude oil gathering and transmission pipelines, more than 350,000 b/d of operating capacity, about 1.2 million bbl of operational storage capacity, and Silver Creek’s 49% non-operated interest in the Tallgrass Energy-led Powder River Gateway Joint Venture, which owns and operates both the Iron Horse and Powder River Express pipelines. Plains said the acquisition will enhance producer access to its integrated transportation network, including gathering and transmission systems in the Powder River Basin and long-haul pipelines delivering crude oil to Cushing. The company also said the assets are supported by a diversified customer base, 915,000 dedicated acres under long-term acreage dedications and minimum volume commitments, and contracts with a weighted-average remaining term of more than eight years. Current throughput averages about 125,000 b/d. The transaction is expected to close in this year’s

Read More »

From Announcements to Delivery: What Separates Real AI Data Center Projects From the Rest

The AI infrastructure market has become very good at announcing gigawatts. Delivering them is another matter. That distinction framed one of the closing sessions of Day 1 at the Data Center Frontier Trends Summit 2026 (Aug. 4-6), where Sean Farney, vice president of data center strategy at JLL and a member of the Data Center Frontier Editorial Advisory Board, moderated a discussion on why some AI data center projects advance from concept to construction while others remain little more than ambitious site plans. Farney was joined by Lawrence Vo, vice president of M&A and capex at Csquare; John Day, chief commercial officer at CleanArc Data Centers; Justin Loth, executive director of power development at Provident Data Centers; and Roshan Shah, co-founder and CEO of Decimal Digital. The question Farney put before the group was straightforward: amid a market moving at what he called “the speed of light,” what separates the developers that actually get projects done from those that do not? The answers repeatedly came back to the same point. In the current market, land, capital and an announcement are no longer enough. Developers have to prove that power is deliverable, infrastructure is ready, regulatory processes are moving, communities are receptive, talent is available and the commercial model can withstand changing conditions. A Gigawatt on Paper Is Not a Gigawatt of Capacity For Loth, who spent roughly 15 years on the utility side before joining Provident, the scale of current data center proposals alone should force the industry to think differently about what constitutes a credible project. Before the hyperscale and AI expansion, he noted, gigawatts were a measure more commonly associated with cities than individual loads. “A 3.5 gigawatt campus,” Loth said, is roughly equivalent to the native load of Austin or San Antonio. That scale makes the distinction

Read More »

The Future of Data Centers: Biomimicry and Community-Centric Design

As a result, Microsoft has said six additional data centers planned in the region are being designed around biomimicry principles rather than treating landscaping as something added after the engineering work is finished. The change, from landscaping as decoration to ecology as a design input, is now being applied elsewhere. There is already a significant US example, set in Mecklenburg County, Virginia, where Microsoft originally announced the Chase City Conservancy in 2022, as part of a data center development south of Chase City. The completed project, which opened in April 2025, protects more than 230 acres from development. It includes more than eight acres of wetlands, over 16,300 linear feet of restored streams, 185 acres of native pollinator habitat, more than 25,000 planted trees and over three miles of publicly accessible walking trails. Local environmental organizations helped shift the design away from what the company describes as a more conventional recreational area toward biodiversity and habitat conservation illustrating the community-engagement side of Microsoft’s model, which, given the current temperature of such relationships, can’t be understated. For data center developers, that may be as important as the ecological results. Community impact is no longer being evaluated on just tax revenue and jobs. Turning portions of a site into protected wetlands, forests, trails or habitat potentially creates a visible local benefit in ways that renewable-energy contracts hundreds of miles away cannot. Microsoft’s commitment to the local community has been led by their Community First AI Infrastructure Plan announced in January 2026. Wetlands in Wisconsin, Screening in Georgia At Microsoft’s massive Mount Pleasant, Wisconsin, AI data center development, the company is working with the Root-Pike Watershed Initiative Network on restoration projects involving wetlands, native prairie and forested riparian buffers. One element involves returning previously straightened streams to more natural, winding channels, improving aquatic

Read More »

Axelera Europa targets enterprise data centers with far more efficient AI

Software is still the gatekeeper Axelera In terms of software enablement, Axelera’s Voyager SDK spans its existing Metis products and the new Europa architecture, providing a common environment across embedded, edge and server deployments, with support for a multitude of computer vision models, LLMs, VLMs, diffusion models, speech and other AI workloads. To automate setup, Axelera’s Voyager Wingman uses natural-language prompts to help developers create or port inference pipelines, while AxeleraScript, or AxScript, provides a Python-enabled domain-specific language with lower-level AIPU control for custom operators and transformer models. This could prove every bit as important as Europa’s performance and efficiency. Enterprises already have models, development environments and application stacks. Extensive rewriting or specialized expertise adds development and operational costs that can quickly undermine savings on hardware and power.

Read More »

Scott Bergs, CEO of Kirkwood IG: Fiber and the AI Data Center Buildout

For years, fiber was one of the more forgiving elements of data center site selection. Developers could secure land, line up power, begin planning the facility and then work with carriers to establish the connectivity required by tenants. In a traditional multi-tenant data center, that model generally worked. At AI scale, Scott Bergs says it increasingly does not. “The architecture of those original communications service provider networks just don’t meet the latency and/or capacity needs” of today’s high-density compute environments, said Bergs, CEO of Kirkwood Infrastructure Group, during a recent episode of the Data Center Frontier Show. The result is a significant change in the data center development stack: network infrastructure can no longer be treated as something that gets solved after the site is chosen. For hyperscalers and neo-cloud providers, fiber route diversity, latency, physical security and future capacity increasingly need to enter the conversation alongside power and land. And as data center campuses follow available power farther from established digital infrastructure hubs, the scale of the network challenge is expanding with them. A connection between data center campuses that might once have extended two or 30 miles can now stretch 250 miles or more, Bergs said. What would traditionally have been considered a long-haul fiber route is increasingly becoming another piece of inter-campus infrastructure. That change is helping drive Kirkwood’s own expansion. From DF&I to Kirkwood Bergs previously led DF&I, a dark-fiber infrastructure platform concentrated in Northern Virginia and Maryland. Kirkwood Infrastructure Group is not simply DF&I under a new name, he said. Rather, it represents what Bergs described as a second phase in a broader infrastructure investment strategy developed originally through IPI Partners. IPI, an investment platform focused on digital infrastructure, backed DF&I after identifying communications infrastructure serving dense compute environments as an area requiring greater direct

Read More »

Cisco brings Splunk AI on premises, expands agent observability, monitors token costs

Cisco executives said during a press briefing that AI agents are operating across data centers, campuses, and branches, and interacting with enterprise resources and other agents. For example, if an agent deletes thousands of files, IT teams need to determine whether it malfunctioned or was compromised. On-premises option: Cisco AI POD for Splunk Aimed at enterprises that need to keep sensitive machine data within their own environments, Cisco AI POD for Splunk “brings Splunk AI to on-premises customers with new AI runtime software, Cisco infrastructure, Nvidia accelerated computing, and Kubernetes-based architecture, pre-validated and optimized for Splunk AI workloads,” according to Cisco. “One of the biggest roadblocks to enterprise AI today is that it’s too hard to deploy,” said Jeetu Patel, Cisco’s president and chief product officer, in a statement. “Customers want to know: Can I trust it to do the job? Can I afford it? And, most importantly, can I secure it? By running Splunk AI on the infrastructure customers already trust, they can move faster to put AI to work in their business with confidence and control.”

Read More »

Nuclear’s Next AI Test: Building at Scale

For data center developers facing multiyear utility interconnection queues and tightening power markets, nuclear energy is entering a different phase of its AI infrastructure story. The near-term opportunity still rests largely with the existing reactor fleet. Holtec International has moved the Palisades Nuclear Plant in Michigan into fuel loading, one of the final major stages before reactor startup activities. Constellation Energy, meanwhile, continues to work toward a 2027 restart of the former Three Mile Island Unit 1, now the Christopher M. Crane Clean Energy Center, under its long-term power agreement with Microsoft. Together, Palisades and Crane represent roughly 1.64 GW of existing nuclear capacity that could return to service without waiting for entirely new plants to be licensed, financed and constructed. That makes reactor restarts one of the few ways nuclear generation can materially intersect with data center power demand before the end of the decade. But the more consequential change may be taking place further upstream. A burst of activity from advanced nuclear developers at the end of August pointed increasingly toward the industrial systems required to move new reactor designs from demonstrations to repeatable infrastructure. X-energy, TerraPower, GE Vernova Hitachi, Oklo, Westinghouse, Kairos Power and others reported progress involving fuel supply, reactor testing, manufacturing, licensing and commercial deployment. None of these advanced reactor projects will solve the industry’s 2027 or 2028 power shortage. That is no longer the most useful test. The more important question is whether advanced nuclear can begin acquiring the characteristics of an industrial supply chain: dependable fuel, standardized manufacturing, repeatable construction, tested reactor systems and enough commercial certainty for large power customers to plan around deployment schedules measured in years rather than speculation. For data center infrastructure, that is the transition worth watching. Palisades Moves From Restoration to Startup The clearest near-term proof point

Read More »

Microsoft will invest $80B in AI data centers in fiscal 2025

And Microsoft isn’t the only one that is ramping up its investments into AI-enabled data centers. Rival cloud service providers are all investing in either upgrading or opening new data centers to capture a larger chunk of business from developers and users of large language models (LLMs).  In a report published in October 2024, Bloomberg Intelligence estimated that demand for generative AI would push Microsoft, AWS, Google, Oracle, Meta, and Apple would between them devote $200 billion to capex in 2025, up from $110 billion in 2023. Microsoft is one of the biggest spenders, followed closely by Google and AWS, Bloomberg Intelligence said. Its estimate of Microsoft’s capital spending on AI, at $62.4 billion for calendar 2025, is lower than Smith’s claim that the company will invest $80 billion in the fiscal year to June 30, 2025. Both figures, though, are way higher than Microsoft’s 2020 capital expenditure of “just” $17.6 billion. The majority of the increased spending is tied to cloud services and the expansion of AI infrastructure needed to provide compute capacity for OpenAI workloads. Separately, last October Amazon CEO Andy Jassy said his company planned total capex spend of $75 billion in 2024 and even more in 2025, with much of it going to AWS, its cloud computing division.

Read More »

John Deere unveils more autonomous farm machines to address skill labor shortage

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Self-driving tractors might be the path to self-driving cars. John Deere has revealed a new line of autonomous machines and tech across agriculture, construction and commercial landscaping. The Moline, Illinois-based John Deere has been in business for 187 years, yet it’s been a regular as a non-tech company showing off technology at the big tech trade show in Las Vegas and is back at CES 2025 with more autonomous tractors and other vehicles. This is not something we usually cover, but John Deere has a lot of data that is interesting in the big picture of tech. The message from the company is that there aren’t enough skilled farm laborers to do the work that its customers need. It’s been a challenge for most of the last two decades, said Jahmy Hindman, CTO at John Deere, in a briefing. Much of the tech will come this fall and after that. He noted that the average farmer in the U.S. is over 58 and works 12 to 18 hours a day to grow food for us. And he said the American Farm Bureau Federation estimates there are roughly 2.4 million farm jobs that need to be filled annually; and the agricultural work force continues to shrink. (This is my hint to the anti-immigration crowd). John Deere’s autonomous 9RX Tractor. Farmers can oversee it using an app. While each of these industries experiences their own set of challenges, a commonality across all is skilled labor availability. In construction, about 80% percent of contractors struggle to find skilled labor. And in commercial landscaping, 86% of landscaping business owners can’t find labor to fill open positions, he said. “They have to figure out how to do

Read More »

2025 playbook for enterprise AI success, from agents to evals

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More 2025 is poised to be a pivotal year for enterprise AI. The past year has seen rapid innovation, and this year will see the same. This has made it more critical than ever to revisit your AI strategy to stay competitive and create value for your customers. From scaling AI agents to optimizing costs, here are the five critical areas enterprises should prioritize for their AI strategy this year. 1. Agents: the next generation of automation AI agents are no longer theoretical. In 2025, they’re indispensable tools for enterprises looking to streamline operations and enhance customer interactions. Unlike traditional software, agents powered by large language models (LLMs) can make nuanced decisions, navigate complex multi-step tasks, and integrate seamlessly with tools and APIs. At the start of 2024, agents were not ready for prime time, making frustrating mistakes like hallucinating URLs. They started getting better as frontier large language models themselves improved. “Let me put it this way,” said Sam Witteveen, cofounder of Red Dragon, a company that develops agents for companies, and that recently reviewed the 48 agents it built last year. “Interestingly, the ones that we built at the start of the year, a lot of those worked way better at the end of the year just because the models got better.” Witteveen shared this in the video podcast we filmed to discuss these five big trends in detail. Models are getting better and hallucinating less, and they’re also being trained to do agentic tasks. Another feature that the model providers are researching is a way to use the LLM as a judge, and as models get cheaper (something we’ll cover below), companies can use three or more models to

Read More »

OpenAI’s red teaming innovations define new essentials for security leaders in the AI era

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More OpenAI has taken a more aggressive approach to red teaming than its AI competitors, demonstrating its security teams’ advanced capabilities in two areas: multi-step reinforcement and external red teaming. OpenAI recently released two papers that set a new competitive standard for improving the quality, reliability and safety of AI models in these two techniques and more. The first paper, “OpenAI’s Approach to External Red Teaming for AI Models and Systems,” reports that specialized teams outside the company have proven effective in uncovering vulnerabilities that might otherwise have made it into a released model because in-house testing techniques may have missed them. In the second paper, “Diverse and Effective Red Teaming with Auto-Generated Rewards and Multi-Step Reinforcement Learning,” OpenAI introduces an automated framework that relies on iterative reinforcement learning to generate a broad spectrum of novel, wide-ranging attacks. Going all-in on red teaming pays practical, competitive dividends It’s encouraging to see competitive intensity in red teaming growing among AI companies. When Anthropic released its AI red team guidelines in June of last year, it joined AI providers including Google, Microsoft, Nvidia, OpenAI, and even the U.S.’s National Institute of Standards and Technology (NIST), which all had released red teaming frameworks. Investing heavily in red teaming yields tangible benefits for security leaders in any organization. OpenAI’s paper on external red teaming provides a detailed analysis of how the company strives to create specialized external teams that include cybersecurity and subject matter experts. The goal is to see if knowledgeable external teams can defeat models’ security perimeters and find gaps in their security, biases and controls that prompt-based testing couldn’t find. What makes OpenAI’s recent papers noteworthy is how well they define using human-in-the-middle

Read More »