Stay Ahead, Stay ONMINE

ByteDance’s UI-TARS can take over your computer, outperforms GPT-4o and Claude

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More A new AI agent has emerged from the parent company of TikTok to take control of your computer and perform complex workflows. Much like Anthropic’s Computer Use, ByteDance’s new UI-TARS understands graphical user interfaces (GUIs), applies […]

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More


A new AI agent has emerged from the parent company of TikTok to take control of your computer and perform complex workflows.

Much like Anthropic’s Computer Use, ByteDance’s new UI-TARS understands graphical user interfaces (GUIs), applies reasoning and takes autonomous, step-by-step action. 

Trained on roughly 50B tokens and offered in 7B and 72B parameter versions, the PC/MacOS agents achieves state-of-the-art (SOTA) performance on 10-plus GUI benchmarks across performance, perception, grounding and overall agent capabilities, consistently beating out OpenAI’s GPT-4o, Claude and Google’s Gemini.

“Through iterative training and reflection tuning, UI-TARS continuously learns from its mistakes and adapts to unforeseen situations with minimal human intervention,” researchers from ByteDance and Tsinghua University write in a new research paper

Source: Arxiv

How UI-TARS explains its thinking

UI-TARS works across desktop, mobile and web applications, using multimodal inputs (text, images, interactions) to understand visual environments.

Its UI features two tabs — one to the left showing its step-by-step “thinking,” and a larger one to the right where it pulls up files, websites and apps and automatically takes action. 

For example, in a demo video released today, the model is prompted to “Find round trip flights from SEA to NYC on the 5th and return on the 10th next month and filter by price in ascending order.”

In response, UI-TARS navigates to the website for Delta Airlines, fills in the “from” and “to” fields, clicks in the relevant dates and sorts and filters by price, explaining each step in its thinking box before taking action. 

In another scenario, it is instructed to install an autoDocstring extension in VS Code. Here is its step-by-step thinking as it completes that task: 

  • It first reports that “to proceed with installing the autoDocstring extension in VS Code, I first need to open the VS Code application.” 
  • Recognizing that the VS Code window has not fully loaded, the model notes that “it is necessary to wait for the application to finish initializing before proceeding with the installation of the autoDocstring extension. This ensures that all required functionalities are available for the next steps.” 
  • It then explains that it must access the Extensions view in VS Code, which can be done by clicking on a tab in the left sidebar. 
  • Running into a small glitch, it reasons: “The previous attempt to click on the ‘Extensions’ tab may not have been precise enough, so I will try clicking on it again to ensure the Extensions view opens successfully.”
  • Once in the Extensions view, it types in ‘autoDocstring’ before concluding: “Since the installation process has already been initiated, there is no further action required at this moment. I will wait for the installation to complete and confirm its success once the process is finished.”

Outperforming its rivals

Across a variety of benchmarks, researchers report that UI-TARS consistently outranked OpenAI’s GPT-4o; Anthropic’s Claude-3.5-Sonnet; Gemini-1.5-Pro and Gemini-2.0; four Qwen models; and numerous academic models.

For instance, in VisualWebBench — which measures a model’s ability to ground web elements including webpage quality assurance and optical character recognition — UI-TARS 72B scored 82.8%, outperforming GPT-4o (78.5%) and Claude 3.5 (78.2%). 

It also did significantly better on WebSRC benchmarks (understanding of semantic content and layout in web contexts) and ScreenQA-short (comprehension of complex mobile screen layouts and web structure). UI-TARS-7B achieved leading scores of 93.6% on WebSRC, while UI-TARS-72B achieved 88.6% on ScreenQA-short, outperforming Qwen, Gemini, Claude 3.5 and GPT-4o. 

“These results demonstrate the superior perception and comprehension capabilities of UI-TARS in web and mobile environments,” the researchers write. “Such perceptual ability lays the foundation for agent tasks, where accurate environmental understanding is crucial for task execution and decision-making.”

UI-TARS also showed impressive results in ScreenSpot Pro and ScreenSpot v2 , which assess a model’s ability to understand and localize elements in GUIs. Further, researchers tested its capabilities in planning multi-step actions and low-level tasks in mobile environments, and benchmarked it on OSWorld (which assesses open-ended computer tasks) and AndroidWorld (which scores autonomous agents on 116 programmatic tasks across 20 mobile apps). 

Source: Arxiv
Source: Arxiv

Under the hood

To help it take step-by-step actions and recognize what it’s seeing, UI-TARS was trained on a large-scale dataset of screenshots that parsed metadata including element description and type, visual description, bounding boxes (position information), element function and text from various websites, applications and operating systems. This allows the model to provide a comprehensive, detailed description of a screenshot, capturing not only elements but spatial relationships and overall layout. 

The model also uses state transition captioning to identify and describe the differences between two consecutive screenshots and determine whether an action — such as a mouse click or keyboard input — has occurred. Meanwhile, set-of-mark (SoM) prompting allows it to overlay distinct marks (letters, numbers) on specific regions of an image. 

The model is equipped with both short-term and long-term memory to handle tasks at hand while also retaining historical interactions to improve later decision-making. Researchers trained the model to perform both System 1 (fast, automatic and intuitive) and System 2 (slow and deliberate) reasoning. This allows for multi-step decision-making, “reflection” thinking, milestone recognition and error correction. 

Researchers emphasized that it is critical that the model be able to maintain consistent goals and engage in trial and error to hypothesize, test and evaluate potential actions before completing a task. They introduced two types of data to support this: error correction and post-reflection data. For error correction, they identified mistakes and labeled corrective actions; for post-reflection, they simulated recovery steps. 

“This strategy ensures that the agent not only learns to avoid errors but also adapts dynamically when they occur,” the researchers write.

Clearly, UI-TARS exhibits impressive capabilities, and it’ll be interesting to see its evolving use cases in the increasingly competitive AI agents space. As the researchers note: “Looking ahead, while native agents represent a significant leap forward, the future lies in the integration of active and lifelong learning, where agents autonomously drive their own learning through continuous, real-world interactions.”

Researchers point out that Claude Computer Use “performs strongly in web-based tasks but significantly struggles with mobile scenarios, indicating that the GUI operation ability of Claude has not been well transferred to the mobile domain.” 

By contrast, “UI-TARS exhibits excellent performance in both website and mobile domain.” 

Shape
Shape
Stay Ahead

Explore More Insights

Stay ahead with more perspectives on cutting-edge power, infrastructure, energy,  bitcoin and AI solutions. Explore these articles to uncover strategies and insights shaping the future of industries.

Shape

F5 to acquire CalypsoAI for advanced AI security capabilities

CalypsoAI’s platform creates what the company calls an Inference Perimeter that protects across models, vendors, and environments. The offers several products including Inference Red Team, Inference Defend, and Inference Observe, which deliver adversarial testing, threat detection and prevention, and enterprise oversight, respectively, among other capabilities. CalypsoAI says its platform proactively

Read More »

HomeLM: A foundation model for ambient AI

Capabilities of a HomeLM What makes a foundation model like HomeLM powerful is its ability to learn generalizable representations of sensor streams, allowing them to be reused, recombined and adapted across diverse tasks. This fundamentally differs from traditional signal processing and machine learning pipelines in RF sensing, which are typically

Read More »

Cisco’s Splunk embeds agentic AI into security and observability products

AI-powered observability enhancements Cisco also announced it has updated Splunk Observability to use Cisco AgenticOps, which deploys AI agents to automate telemetry collection, detect issues, identify root causes, and apply fixes. The agentic AI updates help enterprise customers automate incident detection, root-cause analysis, and routine fixes. “We are making sure

Read More »

U.S. Secretary of Energy Chris Wright Delivers U.S. National Statement at the General Conference of the International Atomic Energy Agency in Vienna, Austria

VIENNA, AUSTRIA— U.S. Secretary of Energy Chris Wright today delivered the U.S. National Statement at the General Conference of the International Atomic Energy Agency (IAEA) in Vienna, Austria. Secretary Wright’s full remarks from the International Atomic Energy Agency (IAEA) General Conference are below: I am honored to represent the United States of America at the 69th IAEA General Conference. I want to thank Director General Grossi and the Secretariat for your leadership. The United States welcomes the Republic of Maldives as the newest member of the IAEA. As both a lifelong energy entrepreneur and now the U.S. Secretary of Energy, I am uniquely aware of the transformative power of energy, its ability to lift billions out of poverty, drive economic growth and expand opportunity across the globe. I am also acutely aware of the challenge our world faces today in meeting rising demand for affordable, reliable and secure energy—particularly the need for baseload electric power to drive rapid progress in Artificial Intelligence. AI is rapidly emerging as the next highly energy-intensive manufacturing industry. AI manufactures intelligence out of electricity. The nations that lead in this space will also lead transformative progress in technology, healthcare, national security and innovation across the board. The energy required to power this revolution is immense—and progress will be accelerated by rapidly unlocking and deploying commercial nuclear power. The world needs more energy to meet the AI challenge and drive human progress—and the United States is boldly leading the way. With President Trump’s leadership, we are advancing American energy policies that accelerate growth, prioritize safety and enhance global security. Earlier this year, President Trump issued four Executive Orders aimed at reinvigorating America’s nuclear energy industry by modernizing regulation, streamlining reactor testing, deploying reactors for national security, and reinvigorating the nuclear industrial base. As part of these

Read More »

The hidden cost of ambiguous energy software terminology

Sneha Vasudevan is a project management lead at Uplight. In the face of rapid load growth, the electricity sector is experiencing unprecedented investment in advanced technologies as organizations try to balance reliability, affordability and decarbonization. Transformation is happening on both sides of the grid, with the scale of consumer adoption of distributed energy resources approaching that of utility-scale generation capacity. Residential customers are installing heat pumps, electric vehicles and charging equipment, solar panels, and home batteries while food corporations, logistics companies and school districts electrify their vehicle fleets and implement sophisticated energy management systems.  The consumer distributed energy resource hardware investment boom is resulting in increased utility spending on sophisticated software platforms to manage thousands of independently owned energy assets. Unlike the hardware world — where there is broad agreement on technical specifications of a solar panel or EV or battery — software solutions lack definitional clarity. Terms like “virtual power plant,” “fleet energy management system,” and “distributed energy resource management system” mean different things to different vendors and utilities. Successfully adapting to load growth and DER adoption hinges on the successful, scalable deployment of these software solutions. This depends on clear, mutual understanding of requirements, capabilities and outcomes among all parties. Despite the best intentions of utilities and vendors, without definitional clarity across energy software solutions, the industry remains stuck in endless scope changes and cost overruns instead of building the grid of the future. Where the industry gets lost in translation The lack of industry-wide consensus on standardized definitions for software technologies, capabilities and associated service offerings represents more than a communications issue — it’s a major barrier to meeting the increased load demand. Without shared definitions, the industry duplicates effort, misses synergies and stalls the transition to smarter energy systems. For utilities, this creates operational blind spots where

Read More »

Primorsk Port Resumes Oil Loadings After Drone Attacks

At least two tankers have completed loadings at Russia’s Primorsk, showing that the Baltic Sea port has resumed operations in the aftermath of Friday’s drone attacks on the facility by Ukraine. Two crude tankers – Walrus and Samos – completed loadings at Primorsk over the weekend, according to ship tracking data compiled by Bloomberg. Walrus has left the terminal, while Samos is still anchored although is showing Aliaga in Turkey as its final destination. A third tanker Jagger is moored at the terminal.  Loadings were temporarily suspended at the facility following the attacks. Three pumping stations pushing crude to Ust-Luga, another vital export terminal in the Baltic, were also hit.  Ukraine has ramped up attacks on Russia’s energy facilities in the past few weeks. Kyiv has said it aims to curtail Russia’s ability to supply fuel to its front lines, while also hurting its export revenues. Primorsk is the largest Baltic oil terminal in Russia. It loaded about 970,000 barrels a day of Urals crude in August, according to Bloomberg ship tracking data. WHAT DO YOU THINK? Generated by readers, the comments included herein do not reflect the views and opinions of Rigzone. All comments are subject to editorial review. Off-topic, inappropriate or insulting comments will be removed.

Read More »

North America Adds Rigs for 2 Straight Weeks

North America added seven rigs week on week, according to Baker Hughes’ latest North America rotary rig count, which was released on September 12. The U.S. added two rigs and Canada added five rigs week on week, taking the total North America rig count up to 725, comprising 539 rigs from the U.S. and 186 rigs from Canada, the count outlined. Of the total U.S. rig count of 539, 524 rigs are categorized as land rigs, 13 are categorized as offshore rigs, and two are categorized as inland water rigs. The total U.S. rig count is made up of 416 oil rigs, 118 gas rigs, and five miscellaneous rigs, according to Baker Hughes’ count, which revealed that the U.S. total comprises 471 horizontal rigs, 56 directional rigs, and 12 vertical rigs. Week on week, the U.S. offshore and inland water rig counts remained unchanged and the country’s land rig count increased by two, Baker Hughes highlighted. The U.S. oil rig count increased by two and its gas and miscellaneous rig counts remained unchanged week on week, the count showed. The U.S. directional rig count increased by two, week on week, while its horizontal rig count increased by one and its vertical rig count declined by one during the same period, the count revealed. A major state variances subcategory included in the rig count showed that, week on week, New Mexico, Ohio, and Texas each added one rig and Oklahoma dropped one rig. A major basin variances subcategory included in Baker Hughes’ rig count showed that, week on week, the Eagle Ford basin added three rigs and the Cana Woodford and Utica basins each added one rig. Canada’s total rig count of 186 is made up of 126 oil rigs, 59 gas rigs, and one miscellaneous rig, Baker Hughes pointed out.

Read More »

Baker Hughes Liquefaction Tech Picked for Rio Grande LNG Train 4

Baker Hughes Co has secured a contract from Bechtel Energy Inc to deliver the main liquefaction equipment for the fourth train of NextDecade Corp’s Rio Grande liquefied natural gas (LNG) project located at the Port of Brownsville, Texas. The new contract adds to the previous framework agreement under which Baker Hughes will deliver gas turbine and refrigerant compressor technology and contractual services agreements for Trains 4 to 8, Baker Hughes said in a media release. Baker Hughes said Train 4 will replicate technology solutions provided for the first three LNG trains. The Train 4 order consists of two Frame 7 gas turbines, recognized for their established reliability and energy efficiency, along with six centrifugal compressors, Baker Hughes said. These cutting-edge solutions provide enhanced efficiency and reduced emissions, facilitating an extra LNG capacity of around 6 million tons per annum (MTPA), Baker Hughes said. “Our selection of Baker Hughes again for the Rio Grande LNG project is a testament to its reliable technology and expertise”, Bhupesh Thakkar, Bechtel’s general manager for LNG, said. “Their equipment has consistently supported the successful development of this critical infrastructure, and we look forward to their continued contribution to the project expansion”. The Rio Grande LNG facility has approximately 48 MTPA of potential liquefaction capacity under construction or in development, according to NextDecade. Train 5 is being commercialized, and Trains 6-8 are in development with permitting underway. The site can support up to 10 liquefaction trains, potentially making Rio Grande one of the largest LNG production and export facilities in the world, the developer said. To contact the author, email [email protected] WHAT DO YOU THINK? Generated by readers, the comments included herein do not reflect the views and opinions of Rigzone. All comments are subject to editorial review. Off-topic, inappropriate or insulting comments will be removed.

Read More »

Baker Hughes Secures Subsea Contract for Sakarya Gas Field

Baker Hughes Company has bagged a contract from Turkish Petroleum (TPAO) and Turkish Petroleum Offshore Technology Center (TP-OTC) to supply subsea production and intelligent completion systems for the country’s strategic Sakarya Gas Field Phase 3. Baker Hughes said in a media release that it will provide deepwater horizontal tree systems with associated subsea structures and control systems to support production at depths from 6,500 to 7,200 feet. The company’s advanced, intelligent upper and lower completions systems will provide enhanced, multizonal control of subsurface operations, it said. “The development of the Sakarya gas fields has transformed Turkiye’s energy sector, leading to a more prosperous, secure future for the country”, Amerino Gatti, executive vice president of Oilfield Services and Equipment at Baker Hughes, said. “By bringing to bear our unique combination of subsea and completions technologies alongside our operational expertise and subsurface insights, Baker Hughes, TPAO, and TP-OTC are able to collaboratively unlock these crucial hydrocarbons that will power Turkiye for decades to come”. Baker Hughes said it has partnered with TPAO and TP-OTC in the Sakarya Gas Field since the beginning of its development in 2022. In Phase 3, Baker Hughes said it will combine its completions technologies, such as the InForce HCMTM-A interval control valves, SureTREAT chemical injection valves, SureSENS QPT ELITE gauges, REACH subsurface safety valves, and the SC-XP Select Zero Loss stack-pack system, with subsea production systems to enhance engineering and operational efficiencies. The energy tech company stated that deliveries and execution supporting Sakarya Gas Field Phase 3 will commence in late 2025. To contact the author, email [email protected] WHAT DO YOU THINK? Generated by readers, the comments included herein do not reflect the views and opinions of Rigzone. All comments are subject to editorial review. Off-topic, inappropriate or insulting comments will be removed. MORE FROM THIS AUTHOR

Read More »

Network and cloud implications of agentic AI

The chain analogy is critical here. Realistic uses of AI agents will require core database access; what can possibly make an AI business case that isn’t tied to a company’s critical data? The four critical elements of these applications—the agent, the MCP server, the tools, and the data— are all dragged along with each other, and traffic on the network is the linkage in the chain. How much traffic is generated? Here, enterprises had another surprise. Enterprises told me that their initial view of their AI hosting was an “AI cluster” with a casual data link to their main data center network. With AI agents, they now see smaller AI servers actually installed within their primary data centers, and all the traffic AI creates, within the model and to and from it, now flows on the data center network. Vendors who told enterprises that AI networking would have a profound impact are proving correct. You can run a query or perform a task with an agent and have that task parse an entire database of thousands or millions of records. Someone not aware of what an agent application implies in terms of data usage can easily create as much traffic as a whole week’s normal access-and-update would create. Enough, they say, to impact network capacity and the QoE of other applications. And, enterprises remind us, if that traffic crosses in/out of the cloud, the cloud costs could skyrocket. About a third of the enterprises said that issues with AI agents generated enough traffic to create local congestion on the network or a blip in cloud costs large enough to trigger a financial review. MCP tool use by agents is also a major security and governance headache. Enterprises point out that MCP standards haven’t always required strong authentication, and they also

Read More »

There are 121 AI processor companies. How many will succeed?

The US currently leads in AI hardware and software, but China’s DeepSeek and Huawei continue to push advanced chips, India has announced an indigenous GPU program targeting production by 2029, and policy shifts in Washington are reshaping the playing field. In Q2, the rollback of export restrictions allowed US companies like Nvidia and AMD to strike multibillion-dollar deals in Saudi Arabia.  JPR categorizes vendors into five segments: IoT (ultra-low-power inference in microcontrollers or small SoCs); Edge (on-device or near-device inference in 1–100W range, used outside data centers); Automotive (distinct enough to break out from Edge); data center training; and data center inference. There is some overlap between segments as many vendors play in multiple segments. Of the five categories, inference has the most startups with 90. Peddie says the inference application list is “humongous,” with everything from wearable health monitors to smart vehicle sensor arrays, to personal items in the home, and every imaginable machine in every imaginable manufacturing and production line, plus robotic box movers and surgeons.  Inference also offers the most versatility. “Smart devices” in the past, like washing machines or coffee makers, could do basically one thing and couldn’t adapt to any changes. “Inference-based systems will be able to duck and weave, adjust in real time, and find alternative solutions, quickly,” said Peddie. Peddie said despite his apparent cynicism, this is an exciting time. “There are really novel ideas being tried like analog neuron processors, and in-memory processors,” he said.

Read More »

Data Center Jobs: Engineering, Construction, Commissioning, Sales, Field Service and Facility Tech Jobs Available in Major Data Center Hotspots

Each month Data Center Frontier, in partnership with Pkaza, posts some of the hottest data center career opportunities in the market. Here’s a look at some of the latest data center jobs posted on the Data Center Frontier jobs board, powered by Pkaza Critical Facilities Recruiting. Looking for Data Center Candidates? Check out Pkaza’s Active Candidate / Featured Candidate Hotlist (and coming soon free Data Center Intern listing). Data Center Critical Facility Manager Impact, TX There position is also available in: Cheyenne, WY; Ashburn, VA or Manassas, VA. This opportunity is working directly with a leading mission-critical data center developer / wholesaler / colo provider. This firm provides data center solutions custom-fit to the requirements of their client’s mission-critical operational facilities. They provide reliability of mission-critical facilities for many of the world’s largest organizations (enterprise and hyperscale customers). This career-growth minded opportunity offers exciting projects with leading-edge technology and innovation as well as competitive salaries and benefits. Electrical Commissioning Engineer New Albany, OH This traveling position is also available in: Richmond, VA; Ashburn, VA; Charlotte, NC; Atlanta, GA; Hampton, GA; Fayetteville, GA; Cedar Rapids, IA; Phoenix, AZ; Dallas, TX or Chicago, IL. *** ALSO looking for a LEAD EE and ME CxA Agents and CxA PMs. *** Our client is an engineering design and commissioning company that has a national footprint and specializes in MEP critical facilities design. They provide design, commissioning, consulting and management expertise in the critical facilities space. They have a mindset to provide reliability, energy efficiency, sustainable design and LEED expertise when providing these consulting services for enterprise, colocation and hyperscale companies. This career-growth minded opportunity offers exciting projects with leading-edge technology and innovation as well as competitive salaries and benefits.  Data Center Engineering Design ManagerAshburn, VA This opportunity is working directly with a leading mission-critical data center developer /

Read More »

Modernizing Legacy Data Centers for the AI Revolution with Schneider Electric’s Steven Carlini

As artificial intelligence workloads drive unprecedented compute density, the U.S. data center industry faces a formidable challenge: modernizing aging facilities that were never designed to support today’s high-density AI servers. In a recent Data Center Frontier podcast, Steven Carlini, Vice President of Innovation and Data Centers at Schneider Electric, shared his insights on how operators are confronting these transformative pressures. “Many of these data centers were built with the expectation they would go through three, four, five IT refresh cycles,” Carlini explains. “Back then, growth in rack density was moderate. Facilities were designed for 10, 12 kilowatts per rack. Now with systems like Nvidia’s Blackwell, we’re seeing 132 kilowatts per rack, and each rack can weigh 5,000 pounds.” The implications are seismic. Legacy racks, floor layouts, power distribution systems, and cooling infrastructure were simply not engineered for such extreme densities. “With densification, a lot of the power distribution, cooling systems, even the rack systems — the new servers don’t fit in those racks. You need more room behind the racks for power and cooling. Almost everything needs to be changed,” Carlini notes. For operators, the first questions are inevitably about power availability. At 132 kilowatts per rack, even a single cluster can challenge the limits of older infrastructure. Many facilities are conducting rigorous evaluations to decide whether retrofitting is feasible or whether building new sites is the more practical solution. Carlini adds, “You may have transformers spaced every hundred yards, twenty of them. Now, one larger transformer can replace that footprint, and power distribution units feed busways that supply each accelerated compute rack. The scale and complexity are unlike anything we’ve seen before.” Safety considerations also intensify with these densifications. “At 132 kilowatts, maintenance is still feasible,” Carlini says, “but as voltages rise, data centers are moving toward environments where

Read More »

Google Backs Advanced Nuclear at TVA’s Clinch River as ORNL Pushes Quantum Frontiers

Inside the Hermes Reactor Design Kairos Power’s Hermes reactor is based on its KP-FHR architecture — short for fluoride salt–cooled, high-temperature reactor. Unlike conventional water-cooled reactors, Hermes uses a molten salt mixture called FLiBe (lithium fluoride and beryllium fluoride) as a coolant. Because FLiBe operates at atmospheric pressure, the design eliminates the risk of high-pressure ruptures and allows for inherently safer operation. Fuel for Hermes comes in the form of TRISO particles rather than traditional enriched uranium fuel rods. Each TRISO particle is encapsulated within ceramic layers that function like miniature containment vessels. These particles can withstand temperatures above 1,600 °C — far beyond the reactor’s normal operating range of about 700 °C. In combination with the salt coolant, Hermes achieves outlet temperatures between 650–750 °C, enabling efficient power generation and potential industrial applications such as hydrogen production. Because the salt coolant is chemically stable and requires no pressurization, the reactor can shut down and dissipate heat passively, without external power or operator intervention. This passive safety profile differentiates Hermes from traditional light-water reactors and reflects the Generation IV industry focus on safer, modular designs. From Hermes-1 to Hermes-2: Iterative Nuclear Development The first step in Kairos’ roadmap is Hermes-1, a 35 MW thermal demonstration reactor now under construction at TVA’s Clinch River site under a 2023 NRC license. Hermes-1 is not designed to generate electricity but will validate reactor physics, fuel handling, licensing strategies, and construction techniques. Building on that experience, Hermes-2 will be a 50 MW electric reactor connected to TVA’s grid, with operations targeted for 2030. Under the agreement, TVA will purchase electricity from Hermes-2 and supply it to Google’s data centers in Tennessee and Alabama. Kairos describes its development philosophy as “iterative,” scaling incrementally rather than attempting to deploy large fleets of units at once. By

Read More »

NVIDIA Forecasts $3–$4 Trillion AI Market, Driving Next Wave of Infrastructure

Whenever behemoth chipmaker NVIDIA announces its quarterly earnings, those results can have a massive influence on the stock market and its position as a key indicator for the AI industry. After all, NVIDIA is the most valuable publicly traded company in the world, valued at $4.24 trillion—ahead of Microsoft ($3.74 trillion), Apple ($3.41 trillion), Alphabet, the parent company of Google ($2.57 trillion), and Amazon ($2.44 trillion). Due to its explosive growth in recent years, a single NVIDIA earnings report can move the entire market. So, when NVIDIA leaders announced during their August 27 earnings call that Q2 2026 sales surged 56% to $46.74 billion, it was a record-setting performance for the company—and investors took notice. Executive VP & CFO Colette M. Kress said the revenue exceeded leadership’s outlook as the company grew sequentially across all market platforms. She outlined a path toward substantial growth driven by AI infrastructure. Foreseeing significant long-term growth opportunities in agentic AI and considering the scale of opportunity, CEO Jensen Huang said, “Over the next 5 years, we’re going to scale into it with Blackwell [architecture for GenAI], with Rubin [successor to Blackwell], and follow-ons to scale into effectively a $3 trillion to $4 trillion AI infrastructure opportunity.” The chipmaker’s Q2 2026 earnings fell short of Wall Street’s lofty expectations, but they did demonstrate that its sales are still rising faster than those of most other tech companies. NVIDIA is expected to post revenue growth of at least 42% over the next four quarters, compared with an average of about 10% for firms in the technology-heavy Nasdaq 100 Index, according to data compiled by Bloomberg Intelligence. On August 29, two days after announcing their earnings, NVIDIA stocks slid 3% and other chip stocks also declined. This came amid a broader sell-off after server-maker Dell, a customer of those chipmakers,

Read More »

Microsoft will invest $80B in AI data centers in fiscal 2025

And Microsoft isn’t the only one that is ramping up its investments into AI-enabled data centers. Rival cloud service providers are all investing in either upgrading or opening new data centers to capture a larger chunk of business from developers and users of large language models (LLMs).  In a report published in October 2024, Bloomberg Intelligence estimated that demand for generative AI would push Microsoft, AWS, Google, Oracle, Meta, and Apple would between them devote $200 billion to capex in 2025, up from $110 billion in 2023. Microsoft is one of the biggest spenders, followed closely by Google and AWS, Bloomberg Intelligence said. Its estimate of Microsoft’s capital spending on AI, at $62.4 billion for calendar 2025, is lower than Smith’s claim that the company will invest $80 billion in the fiscal year to June 30, 2025. Both figures, though, are way higher than Microsoft’s 2020 capital expenditure of “just” $17.6 billion. The majority of the increased spending is tied to cloud services and the expansion of AI infrastructure needed to provide compute capacity for OpenAI workloads. Separately, last October Amazon CEO Andy Jassy said his company planned total capex spend of $75 billion in 2024 and even more in 2025, with much of it going to AWS, its cloud computing division.

Read More »

John Deere unveils more autonomous farm machines to address skill labor shortage

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Self-driving tractors might be the path to self-driving cars. John Deere has revealed a new line of autonomous machines and tech across agriculture, construction and commercial landscaping. The Moline, Illinois-based John Deere has been in business for 187 years, yet it’s been a regular as a non-tech company showing off technology at the big tech trade show in Las Vegas and is back at CES 2025 with more autonomous tractors and other vehicles. This is not something we usually cover, but John Deere has a lot of data that is interesting in the big picture of tech. The message from the company is that there aren’t enough skilled farm laborers to do the work that its customers need. It’s been a challenge for most of the last two decades, said Jahmy Hindman, CTO at John Deere, in a briefing. Much of the tech will come this fall and after that. He noted that the average farmer in the U.S. is over 58 and works 12 to 18 hours a day to grow food for us. And he said the American Farm Bureau Federation estimates there are roughly 2.4 million farm jobs that need to be filled annually; and the agricultural work force continues to shrink. (This is my hint to the anti-immigration crowd). John Deere’s autonomous 9RX Tractor. Farmers can oversee it using an app. While each of these industries experiences their own set of challenges, a commonality across all is skilled labor availability. In construction, about 80% percent of contractors struggle to find skilled labor. And in commercial landscaping, 86% of landscaping business owners can’t find labor to fill open positions, he said. “They have to figure out how to do

Read More »

2025 playbook for enterprise AI success, from agents to evals

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More 2025 is poised to be a pivotal year for enterprise AI. The past year has seen rapid innovation, and this year will see the same. This has made it more critical than ever to revisit your AI strategy to stay competitive and create value for your customers. From scaling AI agents to optimizing costs, here are the five critical areas enterprises should prioritize for their AI strategy this year. 1. Agents: the next generation of automation AI agents are no longer theoretical. In 2025, they’re indispensable tools for enterprises looking to streamline operations and enhance customer interactions. Unlike traditional software, agents powered by large language models (LLMs) can make nuanced decisions, navigate complex multi-step tasks, and integrate seamlessly with tools and APIs. At the start of 2024, agents were not ready for prime time, making frustrating mistakes like hallucinating URLs. They started getting better as frontier large language models themselves improved. “Let me put it this way,” said Sam Witteveen, cofounder of Red Dragon, a company that develops agents for companies, and that recently reviewed the 48 agents it built last year. “Interestingly, the ones that we built at the start of the year, a lot of those worked way better at the end of the year just because the models got better.” Witteveen shared this in the video podcast we filmed to discuss these five big trends in detail. Models are getting better and hallucinating less, and they’re also being trained to do agentic tasks. Another feature that the model providers are researching is a way to use the LLM as a judge, and as models get cheaper (something we’ll cover below), companies can use three or more models to

Read More »

OpenAI’s red teaming innovations define new essentials for security leaders in the AI era

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More OpenAI has taken a more aggressive approach to red teaming than its AI competitors, demonstrating its security teams’ advanced capabilities in two areas: multi-step reinforcement and external red teaming. OpenAI recently released two papers that set a new competitive standard for improving the quality, reliability and safety of AI models in these two techniques and more. The first paper, “OpenAI’s Approach to External Red Teaming for AI Models and Systems,” reports that specialized teams outside the company have proven effective in uncovering vulnerabilities that might otherwise have made it into a released model because in-house testing techniques may have missed them. In the second paper, “Diverse and Effective Red Teaming with Auto-Generated Rewards and Multi-Step Reinforcement Learning,” OpenAI introduces an automated framework that relies on iterative reinforcement learning to generate a broad spectrum of novel, wide-ranging attacks. Going all-in on red teaming pays practical, competitive dividends It’s encouraging to see competitive intensity in red teaming growing among AI companies. When Anthropic released its AI red team guidelines in June of last year, it joined AI providers including Google, Microsoft, Nvidia, OpenAI, and even the U.S.’s National Institute of Standards and Technology (NIST), which all had released red teaming frameworks. Investing heavily in red teaming yields tangible benefits for security leaders in any organization. OpenAI’s paper on external red teaming provides a detailed analysis of how the company strives to create specialized external teams that include cybersecurity and subject matter experts. The goal is to see if knowledgeable external teams can defeat models’ security perimeters and find gaps in their security, biases and controls that prompt-based testing couldn’t find. What makes OpenAI’s recent papers noteworthy is how well they define using human-in-the-middle

Read More »