Stay Ahead, Stay ONMINE

The rise of browser-use agents: Why Convergence’s Proxy is beating OpenAI’s Operator

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More A new wave of AI-powered browser-use agents is emerging, promising to transform how enterprises interact with the web. These agents can autonomously navigate websites, retrieve information, and even complete transactions – but early testing reveals significant […]

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More


A new wave of AI-powered browser-use agents is emerging, promising to transform how enterprises interact with the web. These agents can autonomously navigate websites, retrieve information, and even complete transactions – but early testing reveals significant gaps between promise and performance.

While consumer examples offered by OpenAI’s new browser-use agent Operator, like ordering pizza or buying game tickets, have grabbed headlines, the question is about where the main developer and enterprise use cases are. “The thing that we don’t know is what will be the killer app,” said Sam Witteveen, co-founder of Red Dragon, a company that develops AI agent applications. “My guess is it’s going to be things that just take time on the web that you don’t actually enjoy.” This includes things like going on the web and searching for the cheapest price of a product or booking the best hotel accommodations. More likely it will be used in combination with other tools like Deep Research, where companies can then do even more sophisticated research plus execution of tasks around the web.

Companies need to carefully evaluate the rapidly evolving landscape as established players and startups take different approaches to solving the autonomous browsing challenge.

Key players in the browser-use agent landscape

The field has quickly become crowded with both major tech companies and innovative startups:

Operator and Proxy are the most advanced, in terms of being consumer-friendly and out-of-the-box ready. Many of the others appear to be positioning themselves more for developer or enterprise usage. For example, Browser Use, a Y-Combinator startup that allows users to customize the models used with the agent. This gives you more control over how the agent works, including using a model from your local machine. But it’s definitely more involved.

The others listed above provide a varying degree of functionality and interaction with local machine resources. I decided not even to test ByteDance’s UI-TARS for now, because it requested lower level access to my machine’s security and privacy features (if I test it out, I’ll definitely use a secondary computer). 

Testing reveals reasoning challenges

So the easiest to test are OpenAI’s Operator and Convergence’s Proxy. In our testing, the results highlighted how reasoning capabilities can matter more than raw automation features. Operator, in particular, was more buggy.

For example, I asked the agents to find and summarize VentureBeat’s five most popular stories. It was an ambiguous task, because VentureBeat doesn’t have a “most popular” section per se. Operator struggled with this. It first fell into an infinite scrolling loop while searching for ‘most popular’ stories, requiring manual intervention. In another attempt, it found a three-year-old article titled “Top five stories of the week.” In contrast, Proxy demonstrated better reasoning by identifying the five most visible stories on the homepage as a practical proxy for popularity, and it gave accurate summaries.

The distinction became even clearer in real-world tasks. I asked the agents to book a reservation at a romantic restaurant for noon in Napa, California. Operator approached the task linearly — finding a romantic restaurant first, then checking availability at noon. When no tables were available, it reached a dead end. Proxy showed more sophisticated reasoning by starting with OpenTable to find restaurants that were both romantic and available at the desired time. It even came back with a slightly better rated restaurant.

Even seemingly simple tasks revealed important differences. When searching for a “YubiKey 5C NFC price” on Amazon, Proxy quickly found the item more easily than Operator. 

OpenAI hasn’t divulged much about technologies it uses for training its Operator agent, other than saying it has trained its model on browser-use tasks. Convergence, however, has provided more detail: Its agent uses something called Generative Tree Search to “leverage Web-World Models that predict the state of the web after a proposed action has been taken. These are generated recursively to produce a tree of possible futures that are searched over to select the next optimal action, as ranked by our value models. Our Web-World models can also be used to train agents in hypothetical situations without generating a lot of expensive data.” (More here).

Benchmarks may be useless for now

On paper, these tools appear closely matched. Convergence’s Proxy achieves 88% on the WebVoyager benchmark, which evaluates web agents across 643 real-world tasks on 15 popular websites like Amazon and Booking.com. OpenAI’s Operator scores 87%, while Browser-Use says it reaches 89% but only after changing the WebVoyager codebase slightly, it conceded, “according to our needs”.

These benchmark scores should really be taken with a grain of salt, though, as they can be gamed. The real test comes in practical usage for real-world cases. It’s very early, the space is so rapidly changing, and these products are changing almost on a daily basis. The results will depend more on the specific jobs you’re trying to do, and you may want to instead rely on the vibes you get while using the different products.

Enterprise implications

The implications for enterprise automation are significant. As Witteveen points out in our video podcast conversation about this, where we do a deep dive into this browser-use trend, many companies are currently paying for virtual assistants – operated by real people – to handle basic web research and data gathering tasks. These browser-use agents could dramatically change that equation.

“If AI takes this over,” Witteveen notes, “that’s going to be some of the first low hanging fruit of people losing their jobs. It’s going to show up in some of these kinds of things.”

This could feed into the robotic process automation (RPA) trend, where browser use is pulled in as just another tool for companies to automate more tasks. And as mentioned earlier, the more powerful uses cases will be when an agent combined browser use with other tools, including things like Deep Research, where an LLM-driven agent uses a search tool plus browser use to do more sophisticated jobs.

Cost dynamics driving innovation

Another key factor driving rapid development is the availability of powerful open-source reasoning models like DeepSeek-R1. This allows companies building these browser-use agents to compete effectively with larger players by leveraging these models rather than building their own.

The pricing pressure is already evident. While OpenAI requires a $200 monthly ChatGPT Pro subscription to access Operator, Convergence offers limited free use (up to five uses per day) and a $20/month unlimited plan. This competitive dynamic should accelerate enterprise adoption, though clear use cases are still emerging.

Security and integration challenges

Several hurdles remain before widespread enterprise adoption. Some websites actively block automated browsing, while others require CAPTCHA verification. While OpenAI and Convergence have tools that can get past CAPTCHAs, they let users take over the task to fill them out — instead of doing them directly, since the whole point of CAPTCHAs is to ensure a human is at the other end. Tools like ByteDance’s UI-TARS request deep system access, which raises security concerns for enterprise deployment.

Additionally, the approach to website cooperation varies. OpenAI has worked with specific partners like Instacart, Priceline, DoorDash and Etsy, while others attempt to navigate any website. This inconsistency could impact reliability for enterprise use cases. And of course, any time an agent hits a site requiring login details, that will slow things — as the agents will turn things over to you to fill in those details.

Looking ahead

For enterprises evaluating these tools, the focus should be on specific use cases where autonomous web interaction could provide clear value – whether in research, customer service, or process automation. The technology is progressing rapidly, but success will depend on matching capabilities to concrete business needs.

As this space evolves, expect to see more enterprise-focused features and potentially specialized agents for specific industries or tasks. The race between established players and innovative startups should drive both technical advancement and competitive pricing, making 2025 a crucial year for enterprise browser-use agent adoption.

For more detail on these trends and testing results, check out the full video conversation between Sam Witteveen and myself.

Shape
Shape
Stay Ahead

Explore More Insights

Stay ahead with more perspectives on cutting-edge power, infrastructure, energy,  bitcoin and AI solutions. Explore these articles to uncover strategies and insights shaping the future of industries.

Shape

CoreWeave brings Nvidia Vera Rubin, AI tools to its cloud services

Forge brings together experiment tracking, evaluation, agent observability, post-training, inference and model management. The platform is designed to let teams use their preferred models, frameworks and cloud environments rather than locking them into a single AI stack, CoreWeave stated. The company described Forge as a continuous AI development loop in

Read More »

Memory squeeze set to tighten through 2028, Micron says

Shah said enterprises should lock in pricing too. “Companies should secure multiyear pricing for the computing capacity they know they will need,” he said. Moving workloads to the cloud will not avoid rising hardware and energy costs “because providers will pass them on,” he added. Which refreshes to delay When

Read More »

Cisco SD-WAN Manager hit by zero-day admin access attack

“While every configuration is affected, the practical exposure is not identical across organizations: an internet-accessible management interface presents a much more immediate risk than one isolated within a tightly controlled administrative network,” Grover said, reinforcing Cisco’s advice. Compromising the management layer can give an attacker considerably more leverage than having

Read More »

Energy Department Announces Up to $400 Million for Basic Research to Advance the Frontiers of Science

WASHINGTON—The U.S. Department of Energy (DOE) today announced its annual funding opportunity of up to $400 million to advance the frontiers of scientific knowledge and lay the foundation for future technologies and innovation. The funding delivers on President Trump’s Executive Order Restoring Gold Standard Science by supporting rigorous, transparent, and mission-driven research across DOE’s Office of Science to strengthen America’s scientific and technological leadership. “Foundational research is where the breakthroughs that shape America’s future begin,” said DOE Under Secretary for Science Darío Gil. “Through this investment, we are empowering our nation’s researchers to pursue bold ideas, push the boundaries of discovery, and build the scientific foundations for tomorrow’s technologies—strengthening America’s global leadership in science and innovation.” The funding will support research across DOE’s Office of Science and its major programs to tackle some of the nation’s greatest scientific challenges and accelerate discoveries in areas critical to America’s future. Research will span Advanced Scientific Computing Research, Basic Energy Sciences, Biological and Environmental Research, Fusion Energy Sciences, High Energy Physics, Nuclear Physics, and Isotope R&D and Production. The Notice of Funding Opportunity (NOFO), informally known as the “Open Call,” is issued annually at the beginning of each Fiscal Year (FY). It provides a vehicle for DOE’s Office of Science to solicit applications from institutions for research support in areas not covered by more specific, topical NOFOs issued by the office in FY 2027. More details about this Notice of Funding Opportunity can be found here. DOE’s Office of Science is the nation’s largest supporter of basic research in the physical sciences, funds research at hundreds of universities nationwide, and stewards 10 of DOE’s National Laboratories.

Read More »

Magnolia sells non-core South Texas assets, adds Karnes-area acreage during WildFire integration

The company disclosed the transaction as part of an operational update following the recent closing of its acquisition of WildFire Energy. Chris Stavros, Magnolia chairman, president, and chief executive officer, said integration of the WildFire assets is progressing as planned as the company works to build a larger Eagle Ford and Austin Chalk position across South Texas. Production, capital outlook Third-quarter 2026 production is expected to average 116,000-118,000 boe/d (about 42% oil), reflecting the WildFire acquisition and the impact of the divested properties, Magnolia said. Drilling and completion (D&C) capital spending for the quarter is expected to total $155-165 million. For fourth-quarter 2026, the first full quarter reflecting the WildFire acquisition, production is forecast at 159,000-161,000 boe/d with oil accounting for 49-50% of volumes. D&C spending is expected to be about $235 million. For 2027, Magnolia expects both oil production and total production to grow 4-5% from a second-quarter 2026 pro forma base of about 78,000 bo/d and 158,000 boe/d, respectively, after accounting for volumes associated with the asset sale. The company currently estimates 2027 D&C capital spending of $900-950 million, including the impact of modest oilfield service cost inflation.

Read More »

US Senate negotiators reach permitting bill deal, ending yearslong impasse

US Senate negotiators introduced bipartisan permitting legislation Sept. 30 that would accelerate federal reviews, sharply curtail the window for legal challenges to energy projects, and improve predictability for developers. The agreement is a breakthrough in the Senate, where lawmakers have struggled for years to advance permitting reform. The Bipartisan American Affordability and Jobs Act would impose a 150-day window for certain challenges to federal permitting decisions and provide greater certainty that permitted energy and infrastructure projects retain their approvals. It also would reform Clean Water Act reviews that can delay energy infrastructure, establishing more predictable environmental reviews. The changes could reduce regulatory and litigation uncertainty for interstate gas pipelines, which are often subject to both federal environmental reviews and state water-quality certifications. The four senators behind the deal—Mike Lee (R-Utah), Martin Heinrich (D-NM), Shelley Moore Capito (R-W.Va.), and Sheldon Whitehouse (D-R.I.)—said the agreement also addresses high-profile electricity issues by expanding federal authority over transmission permitting and requiring data centers to pay associated transmission costs rather than shifting them to other electricity customers. The bill, released after months of negotiations, still faces hurdles. Senate negotiators reached an agreement on the legislative text, but Democrats also want assurances that wind and solar projects would benefit from the same streamlining and clarity from the Trump administration on what it means to return wind and solar permitting to “regular order,” issues that remain unresolved. The Senate would consider amendments before voting on the legislation. The target is a vote after the Nov. 3 midterm elections, with Capito noting the permitting bill could be the first vote when the Senate returns. The House has passed its own permitting legislation, the SPEED Act. While both bills would streamline federal permitting and limit litigation, they take different approaches to reforming NEPA, and the Senate legislation addresses transmission

Read More »

MRPL refinery fire at coker-adjacent plant kills one, injures another

Oil & Natural Gas Corp. Ltd. subsidiary Mangalore Refinery and Petrochemicals Ltd. (MRPL) reported one fatality and one injury after a fire at its refinery in Karnataka, Mangalore, India, on Sept. 30. The incident began at about 12:20 p.m. local time when a high-pressure cold separator ruptured in the refinery’s delayed coker gas oil hydrotreater unit, or coker heavy gas oil hydrotreating unit, MRPL said in separate regulatory filings to BSE Ltd. The company isolated the affected unit’s battery limits and deployed emergency response and firefighting teams. MRPL said the fire was brought under control and completely extinguished by 2:47 p.m. local time. During subsequent combing and inspection operations following the fire’s full extinguishing, personnel found the body of one deceased individual at the affected plant. A second person sustained burn injuries and was receiving hospital treatment, the company said. MRPL issued its first update while firefighting operations were continuing. It said one minor injury had been reported at that stage and that the injured person was out of danger. The company did not provide details on the identities of those involved, the cause of the separator rupture, or the extent of damage to the unit. It also did not disclose whether the incident affected broader refinery operations or production. MRPL said further verified information would be released as it becomes available.

Read More »

Dallas Fed survey: More than one in five firms plan to grow capex in 2027

The share of exploration and production (E&P) companies planning to add to their capital spending in 2027 versus this year has grown to 22% from 10% in June, a new Federal Reserve Bank of Dallas survey shows. Of the more than 80 E&P leaders in Texas, northern Louisiana, and southern New Mexico who responded to the latest Dallas Fed Energy Survey earlier this month, a third said their oil production has increased over the past 3 months and only 1 in 8 said they’re pumping less oil. On the capex side, 46% said their spending this quarter was up from this year’s second quarter. Both of those data points were down slightly from the Fed’s June poll. What appears to be changing more substantially on the ground in the Permian basin, Eagle Ford, and other areas in the Dallas Fed’s footprint are expectations about 2027 spending. Only 5% of E&P leaders now expect they’ll trim capex next year while 73% said they’ll keep spending level. Three months ago, those figures were 10% and 81%, respectively. That means 22% of executives now think their capex will climb in 2027 compared to less than 10% 3 months ago. And it suggests that production in the region will climb from here as producers look to take advantage of consistently high prices for their products—even if they’ve retreated from their recent highs. Jon Costello, an analyst at HFI Research, said an industry response—with Texas firms in the vanguard—to higher prices similar to how it recovered starting in late 2016 would grow total US production more than 4% to about 14.4 million b/d.

Read More »

EIA: US crude oil inventories up 900,000 bbl

US crude oil inventories for the week ended Sept. 25, excluding the Strategic Petroleum Reserve, increased by 900,000 bbl from the previous week, according to data from the US Energy Information Administration (EIA). At 427.3 million bbl, US crude oil inventories are 2% above the 5-year average for this time of year, the EIA report indicated. Gasoline output averaged 9.5 million b/d, and distillate production decreased to 5.0 million b/d. Propane-propylene inventories increased 1.8 million bbl, 20% above the 5-year average. Total commercial petroleum inventories decreased by 7 million bbl for the week. Distillate inventories decreased 2.3 million barrels, 14% below the five-year average. US crude oil refinery inputs averaged 16.3 million b/d for the week ended Sept. 25, which was 554,000 b/d less than the previous week’s average. Refineries operated at 92.5% of capacity. Crude oil imports decreased 179,000 million b/d to 5.7 million b/d. The 4-week average of 6.4 million b/d is 4.8% above the year-ago level. Gasoline imports averaged 500,000 b/d; distillate imports averaged 153,000 b/d. Over the past four weeks, total product supplied averaged 20.8 million b/d, up 2.1% year over year. The 4-week average for gasoline product supplied increased 0.3% year over year to 8.7 million b/d, while the 4-week average for distillate product supplied increased 5.2% to 3.8 million b/d. The 4-week average for jet fuel product supplied increased 6.5% year over year.  

Read More »

Data Center Private Power: Who Regulates Behind-the-Meter Generation?

The Cases Writing the Rules The regulatory questions surrounding private power are no longer theoretical. Federal regulators, regional grid operators and state utility commissions are already confronting disputes over co-located generation, transmission obligations, large-load tariffs and the financial commitments required from data center customers. Amazon–Talen Puts Co-Location Before FERC The Amazon–Talen–Susquehanna dispute became the most prominent federal test of how far a co-located data center can separate its power supply from the regional grid while remaining connected to the grid’s reliability framework. Amazon Web Services operates a data center campus adjacent to Talen Energy’s Susquehanna nuclear plant in Pennsylvania. In 2024, the parties sought to expand the co-located load from 300 MW to 480 MW through an amended interconnection agreement with PJM. The proposed structure would have allowed the campus to receive more power directly from the nuclear plant while reducing the plant’s capacity interconnection rights on the PJM system. FERC rejected the amendment in November 2024, concluding that PJM had not justified the nonstandard provisions in the agreement. Importantly, the Commission did not prohibit co-location or dedicated generation. The dispute instead exposed a larger gap in existing transmission rules: how should a large load located beside a generator be measured, what transmission service does it require, what happens when the generator is unavailable, and how should the project pay for continued access to regional reliability? Those questions soon moved beyond the Amazon–Talen project itself. In December 2025, FERC found that PJM’s existing framework did not provide sufficiently clear and consistent rules for generators serving co-located loads or for transmission customers taking service on their behalf. The Commission directed PJM to develop defined interconnection and operating requirements along with transmission options ranging from conventional network service to firm and non-firm contract-demand structures. The central issue is deceptively simple: physical proximity

Read More »

Data Center Thermal Management Series-Part 1 of 3

The biggest driver in today’s global economy is the data center, and in the process, data centers are generating and taking more heat than ever before. Heat — the generation of it and public concerns about it — is the most pressing issue facing the data center industry. In response, data center managers and engineers are viewing the problem in new ways, devising innovative, next-generation thermal control strategies that manage the problem more effectively. In the minds of IT managers and the public, data centers and AI are joined at the hip, along with increasing power demands and associated issues of environmental impact, water usage, and higher utility costs. As community activists, political leaders, and now even some prominent AI industry CEOs call for limits on AI development and data center construction, it’s imperative that the data center industry better manage the heat they’re producing — and taking. Sponsored Resources: Texas Instruments’ portfolio of data center thermal management systems spans the full coolant path. It’s been under active development at TI for decades and has evolved in response to the rising power, computing demands, and complexity of advanced data centers. Traditional thermal management techniques can’t keep up with AI’s intense computing requirements. For decades, air cooling was sufficient. Fans blew cool air across hot components and carried heat away, keeping data centers and AI servers operating reliably. But with AI training and inference pushing rack power beyond 20kW to 40kW, air alone can no longer remove heat quickly enough. In addition, air cooling is noisy and consumes too much energy. AI server racks are coming online now that draw 100kW of power, and they’re on their way to more than a megawatt in a few years. Each generation of servers grows denser, more powerful, and hotter as they move and compute massive volumes

Read More »

LiquidStack unveils liquid-cooling platform targeting AI data centers

“Operators need cooling infrastructure that can adapt as GPU platforms and rack densities evolve,” said Scott Smith, general manager of LiquidStack, in a statement. “CDU 2.X combines the performance and flexibility customers need today with the headroom to prepare for what comes next, allowing them to configure cooling around their facility and deployment strategy rather than designing the facility around the CDU.” A key feature is the platform’s support for different deployment configurations, including end-of-row and rack-adjacent installations. The idea is to give data center operators more flexibility in how they place their equipment with increasingly high thermal loads. The system offers configurable control-valve, power-feed and redundancy options, including dual-feed A/B configurations and automatic transfer switch support. These features allow operators to tailor the CDU to different facility architectures and resiliency requirements.

Read More »

If Apple returns to enterprise server game (with help from Nvidia), what market could it target?

Nvidia introduced NVLink Fusion as part of a broader effort to make its interconnect technology available to companies developing custom AI processors. The approach allows third-party silicon to be incorporated into Nvidia-oriented data center architectures, potentially extending Nvidia’s influence beyond its own GPUs. A similar strategy is already being pursued with inference-chip developer d-Matrix. Apple also already operates specialized servers for its Private Cloud Compute system, which handles Apple Intelligence workloads that require processing beyond the user’s device. Scaling that infrastructure, however, reportedly creates bandwidth, cost, and performance challenges. A commercial AI server would allow Apple to extend its silicon strategy into the data center while giving customers direct access to Apple processors rather than Apple’s own cloud services.

Read More »

From Coal to Compute: How Pennsylvania Is Rebuilding Power for AI Data Centers

Western Pennsylvania is becoming a test bed for one of the most consequential changes underway in data center development: the migration from simply finding grid capacity to building the power supply along with the data center. Two projects illustrate that transition particularly well. Aligned Data Centers is advancing the roughly $10 billion, 2-GW Project Phoenix campus at the former Bruce Mansfield coal-fired power station site in Shippingport, Beaver County. About 60 miles to the east, the former Homer City Generating Station is being transformed into the Homer City Energy Campus, centered on as much as 4.4 GW of new natural-gas generation and a proposed Amazon Web Services campus that could ultimately include 39 data center buildings. The projects share a number of similarities. Both reuse former coal-generation properties. Both already possess much of the infrastructure that greenfield data center developers spend years trying to obtain: high-voltage transmission, industrial zoning, water infrastructure, pipeline access, large parcels and proximity to the Marcellus and Utica natural-gas fields.But technically and commercially, they are also very different. Project Phoenix is essentially a data-center-led development using behind-the-meter generation, supplemented by a separate plan to repower the former coal station. While Homer City is the reverse: a power-generation project being constructed first, with the hyperscale data center development forming around the available power supply. This isn’t a competition, but it is a comparison on speed of delivery and the issues both development models face. Project Phoenix: Aligned Moves Into Pennsylvania Aligned calls Shippingport its first Pennsylvania campus and a regional flagship. The company formally broke ground on Project Phoenix on September 10, 2026, describing it as a 2-GW campus spanning three data center facilities and representing roughly $10 billion in regional investment. Aligned estimates the development will support approximately 3,000 construction jobs and 640 full-time jobs in

Read More »

OpenStack Hibiscus adds DNS security features and confidential computing to the open-source cloud platform

Type-5 support is aimed at data center integration. “It lets the tenant network prefixes be advertised directly into the physical fabric, so that really that’s basically how modern data center networks are built,” Carrez explained. “So really helps OpenStack fit into those environments without extra gateway layers that we’ve seen in use before.” Routable tenant addresses: The OVN BGP integration gains a route leaking option. An operator turns it on with the leak_routes attribute of a subnet. The extended OVN features follow the same approach. “OVN BGP features that let the tenant addresses be routable directly from the underlay, and that again is exposing how modern data centers are built directly into OpenStack,” Carrez said. Lower memory use: In deployments that use Open vSwitch, a monitoring daemon tracked keepalived state changes for HA routers. Hibiscus replaces the daemon with a shell script. The change applies to every HA router, so the savings add up across a deployment. “The 15 times reduction in memory footprint for the high availability router monitoring is, I think, really interesting,” Carrez said.

Read More »

Microsoft will invest $80B in AI data centers in fiscal 2025

And Microsoft isn’t the only one that is ramping up its investments into AI-enabled data centers. Rival cloud service providers are all investing in either upgrading or opening new data centers to capture a larger chunk of business from developers and users of large language models (LLMs).  In a report published in October 2024, Bloomberg Intelligence estimated that demand for generative AI would push Microsoft, AWS, Google, Oracle, Meta, and Apple would between them devote $200 billion to capex in 2025, up from $110 billion in 2023. Microsoft is one of the biggest spenders, followed closely by Google and AWS, Bloomberg Intelligence said. Its estimate of Microsoft’s capital spending on AI, at $62.4 billion for calendar 2025, is lower than Smith’s claim that the company will invest $80 billion in the fiscal year to June 30, 2025. Both figures, though, are way higher than Microsoft’s 2020 capital expenditure of “just” $17.6 billion. The majority of the increased spending is tied to cloud services and the expansion of AI infrastructure needed to provide compute capacity for OpenAI workloads. Separately, last October Amazon CEO Andy Jassy said his company planned total capex spend of $75 billion in 2024 and even more in 2025, with much of it going to AWS, its cloud computing division.

Read More »

John Deere unveils more autonomous farm machines to address skill labor shortage

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Self-driving tractors might be the path to self-driving cars. John Deere has revealed a new line of autonomous machines and tech across agriculture, construction and commercial landscaping. The Moline, Illinois-based John Deere has been in business for 187 years, yet it’s been a regular as a non-tech company showing off technology at the big tech trade show in Las Vegas and is back at CES 2025 with more autonomous tractors and other vehicles. This is not something we usually cover, but John Deere has a lot of data that is interesting in the big picture of tech. The message from the company is that there aren’t enough skilled farm laborers to do the work that its customers need. It’s been a challenge for most of the last two decades, said Jahmy Hindman, CTO at John Deere, in a briefing. Much of the tech will come this fall and after that. He noted that the average farmer in the U.S. is over 58 and works 12 to 18 hours a day to grow food for us. And he said the American Farm Bureau Federation estimates there are roughly 2.4 million farm jobs that need to be filled annually; and the agricultural work force continues to shrink. (This is my hint to the anti-immigration crowd). John Deere’s autonomous 9RX Tractor. Farmers can oversee it using an app. While each of these industries experiences their own set of challenges, a commonality across all is skilled labor availability. In construction, about 80% percent of contractors struggle to find skilled labor. And in commercial landscaping, 86% of landscaping business owners can’t find labor to fill open positions, he said. “They have to figure out how to do

Read More »

2025 playbook for enterprise AI success, from agents to evals

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More 2025 is poised to be a pivotal year for enterprise AI. The past year has seen rapid innovation, and this year will see the same. This has made it more critical than ever to revisit your AI strategy to stay competitive and create value for your customers. From scaling AI agents to optimizing costs, here are the five critical areas enterprises should prioritize for their AI strategy this year. 1. Agents: the next generation of automation AI agents are no longer theoretical. In 2025, they’re indispensable tools for enterprises looking to streamline operations and enhance customer interactions. Unlike traditional software, agents powered by large language models (LLMs) can make nuanced decisions, navigate complex multi-step tasks, and integrate seamlessly with tools and APIs. At the start of 2024, agents were not ready for prime time, making frustrating mistakes like hallucinating URLs. They started getting better as frontier large language models themselves improved. “Let me put it this way,” said Sam Witteveen, cofounder of Red Dragon, a company that develops agents for companies, and that recently reviewed the 48 agents it built last year. “Interestingly, the ones that we built at the start of the year, a lot of those worked way better at the end of the year just because the models got better.” Witteveen shared this in the video podcast we filmed to discuss these five big trends in detail. Models are getting better and hallucinating less, and they’re also being trained to do agentic tasks. Another feature that the model providers are researching is a way to use the LLM as a judge, and as models get cheaper (something we’ll cover below), companies can use three or more models to

Read More »

OpenAI’s red teaming innovations define new essentials for security leaders in the AI era

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More OpenAI has taken a more aggressive approach to red teaming than its AI competitors, demonstrating its security teams’ advanced capabilities in two areas: multi-step reinforcement and external red teaming. OpenAI recently released two papers that set a new competitive standard for improving the quality, reliability and safety of AI models in these two techniques and more. The first paper, “OpenAI’s Approach to External Red Teaming for AI Models and Systems,” reports that specialized teams outside the company have proven effective in uncovering vulnerabilities that might otherwise have made it into a released model because in-house testing techniques may have missed them. In the second paper, “Diverse and Effective Red Teaming with Auto-Generated Rewards and Multi-Step Reinforcement Learning,” OpenAI introduces an automated framework that relies on iterative reinforcement learning to generate a broad spectrum of novel, wide-ranging attacks. Going all-in on red teaming pays practical, competitive dividends It’s encouraging to see competitive intensity in red teaming growing among AI companies. When Anthropic released its AI red team guidelines in June of last year, it joined AI providers including Google, Microsoft, Nvidia, OpenAI, and even the U.S.’s National Institute of Standards and Technology (NIST), which all had released red teaming frameworks. Investing heavily in red teaming yields tangible benefits for security leaders in any organization. OpenAI’s paper on external red teaming provides a detailed analysis of how the company strives to create specialized external teams that include cybersecurity and subject matter experts. The goal is to see if knowledgeable external teams can defeat models’ security perimeters and find gaps in their security, biases and controls that prompt-based testing couldn’t find. What makes OpenAI’s recent papers noteworthy is how well they define using human-in-the-middle

Read More »