Stay Ahead, Stay ONMINE

When AI reasoning goes wrong: Microsoft Research shows more tokens can mean more problems

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Large language models (LLMs) are increasingly capable of complex reasoning through “inference-time scaling,” a set of techniques that allocate more computational resources during inference to generate answers. However, a new study from Microsoft Research reveals that […]

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More


Large language models (LLMs) are increasingly capable of complex reasoning through “inference-time scaling,” a set of techniques that allocate more computational resources during inference to generate answers. However, a new study from Microsoft Research reveals that the effectiveness of these scaling methods isn’t universal. Performance boosts vary significantly across different models, tasks and problem complexities.

The core finding is that simply throwing more compute at a problem during inference doesn’t guarantee better or more efficient results. The findings can help enterprises better understand cost volatility and model reliability as they look to integrate advanced AI reasoning into their applications.

Putting scaling methods to the test

The Microsoft Research team conducted an extensive empirical analysis across nine state-of-the-art foundation models. This included both “conventional” models like GPT-4o, Claude 3.5 Sonnet, Gemini 2.0 Pro and Llama 3.1 405B, as well as models specifically fine-tuned for enhanced reasoning through inference-time scaling. This included OpenAI’s o1 and o3-mini, Anthropic’s Claude 3.7 Sonnet, Google’s Gemini 2 Flash Thinking, and DeepSeek R1.

They evaluated these models using three distinct inference-time scaling approaches:

  1. Standard Chain-of-Thought (CoT): The basic method where the model is prompted to answer step-by-step.
  2. Parallel Scaling: the model generates multiple independent answers for the same question and uses an aggregator (like majority vote or selecting the best-scoring answer) to arrive at a final result.
  3. Sequential Scaling: The model iteratively generates an answer and uses feedback from a critic (potentially from the model itself) to refine the answer in subsequent attempts.

These approaches were tested on eight challenging benchmark datasets covering a wide range of tasks that benefit from step-by-step problem-solving: math and STEM reasoning (AIME, Omni-MATH, GPQA), calendar planning (BA-Calendar), NP-hard problems (3SAT, TSP), navigation (Maze) and spatial reasoning (SpatialMap).

Several benchmarks included problems with varying difficulty levels, allowing for a more nuanced understanding of how scaling behaves as problems become harder.

“The availability of difficulty tags for Omni-MATH, TSP, 3SAT, and BA-Calendar enables us to analyze how accuracy and token usage scale with difficulty in inference-time scaling, which is a perspective that is still underexplored,” the researchers wrote in the paper detailing their findings.

The researchers evaluated the Pareto frontier of LLM reasoning by analyzing both accuracy and the computational cost (i.e., the number of tokens generated). This helps identify how efficiently models achieve their results. 

Inference-time scaling pareto
Inference-time scaling Pareto frontier Credit: arXiv

They also introduced the “conventional-to-reasoning gap” measure, which compares the best possible performance of a conventional model (using an ideal “best-of-N” selection) against the average performance of a reasoning model, estimating the potential gains achievable through better training or verification techniques.

More compute isn’t always the answer

The study provided several crucial insights that challenge common assumptions about inference-time scaling:

Benefits vary significantly: While models tuned for reasoning generally outperform conventional ones on these tasks, the degree of improvement varies greatly depending on the specific domain and task. Gains often diminish as problem complexity increases. For instance, performance improvements seen on math problems didn’t always translate equally to scientific reasoning or planning tasks.

Token inefficiency is rife: The researchers observed high variability in token consumption, even between models achieving similar accuracy. For example, on the AIME 2025 math benchmark, DeepSeek-R1 used over five times more tokens than Claude 3.7 Sonnet for roughly comparable average accuracy. 

More tokens do not lead to higher accuracy: Contrary to the intuitive idea that longer reasoning chains mean better reasoning, the study found this isn’t always true. “Surprisingly, we also observe that longer generations relative to the same model can sometimes be an indicator of models struggling, rather than improved reflection,” the paper states. “Similarly, when comparing different reasoning models, higher token usage is not always associated with better accuracy. These findings motivate the need for more purposeful and cost-effective scaling approaches.”

Cost nondeterminism: Perhaps most concerning for enterprise users, repeated queries to the same model for the same problem can result in highly variable token usage. This means the cost of running a query can fluctuate significantly, even when the model consistently provides the correct answer. 

variance in model outputs
Variance in response length (spikes show smaller variance) Credit: arXiv

The potential in verification mechanisms: Scaling performance consistently improved across all models and benchmarks when simulated with a “perfect verifier” (using the best-of-N results). 

Conventional models sometimes match reasoning models: By significantly increasing inference calls (up to 50x more in some experiments), conventional models like GPT-4o could sometimes approach the performance levels of dedicated reasoning models, particularly on less complex tasks. However, these gains diminished rapidly in highly complex settings, indicating that brute-force scaling has its limits.

GPT-4o inference-time scaling
On some tasks, the accuracy of GPT-4o continues to improve with parallel and sequential scaling. Credit: arXiv

Implications for the enterprise

These findings carry significant weight for developers and enterprise adopters of LLMs. The issue of “cost nondeterminism” is particularly stark and makes budgeting difficult. As the researchers point out, “Ideally, developers and users would prefer models for which the standard deviation on token usage per instance is low for cost predictability.”

“The profiling we do in [the study] could be useful for developers as a tool to pick which models are less volatile for the same prompt or for different prompts,” Besmira Nushi, senior principal research manager at Microsoft Research, told VentureBeat. “Ideally, one would want to pick a model that has low standard deviation for correct inputs.” 

Models that peak blue to the left consistently generate the same number of tokens at the given task Credit: arXiv

The study also provides good insights into the correlation between a model’s accuracy and response length. For example, the following diagram shows that math queries above ~11,000 token length have a very slim chance of being correct, and those generations should either be stopped at that point or restarted with some sequential feedback. However, Nushi points out that models allowing these post hoc mitigations also have a cleaner separation between correct and incorrect samples.

“Ultimately, it is also the responsibility of model builders to think about reducing accuracy and cost non-determinism, and we expect a lot of this to happen as the methods get more mature,” Nushi said. “Alongside cost nondeterminism, accuracy nondeterminism also applies.”

Another important finding is the consistent performance boost from perfect verifiers, which highlights a critical area for future work: building robust and broadly applicable verification mechanisms. 

“The availability of stronger verifiers can have different types of impact,” Nushi said, such as improving foundational training methods for reasoning. “If used efficiently, these can also shorten the reasoning traces.”

Strong verifiers can also become a central part of enterprise agentic AI solutions. Many enterprise stakeholders already have such verifiers in place, which may need to be repurposed for more agentic solutions, such as SAT solvers, logistic validity checkers, etc. 

“The questions for the future are how such existing techniques can be combined with AI-driven interfaces and what is the language that connects the two,” Nushi said. “The necessity of connecting the two comes from the fact that users will not always formulate their queries in a formal way, they will want to use a natural language interface and expect the solutions in a similar format or in a final action (e.g. propose a meeting invite).”

Shape
Shape
Stay Ahead

Explore More Insights

Stay ahead with more perspectives on cutting-edge power, infrastructure, energy,  bitcoin and AI solutions. Explore these articles to uncover strategies and insights shaping the future of industries.

Shape

Cisco rides ‘networking supercycle’ for strong Q4

Security revenue grew 14% year-over-year in Q4, with more than 1,500 customers adopting new products such as Secure Access, XDR, HyperShield, and AI Defense, bringing the total new customer count for these products to 6,400 since launch, Robbins noted. Firewall orders increased more than 30%, and AI security features like AI

Read More »

Equinor lets stimulation service contract for NCS assets

@import url(‘https://fonts.googleapis.com/css2?family=Inter:wght@100..900&display=swap’); .ebm-page__main h1, .ebm-page__main h2, .ebm-page__main h3, .ebm-page__main h4, .ebm-page__main h5, .ebm-page__main h6 { font-family: Inter; } body { line-height: 150%; letter-spacing: 0.025em; } button, .ebm-button-wrapper { font-family: Inter; } .label-style { text-transform: uppercase; color: var(–color-grey); font-weight: 600; font-size: 0.75rem; } .caption-style { font-size: 0.75rem; opacity: .6; } #onetrust-pc-sdk

Read More »

Energy Department Modernizes National Laboratory Operations to Strengthen America’s Scientific, Energy, and National Security Missions

WASHINGTON—The U.S. Department of Energy (DOE) today announced updated operating directives for its National Laboratories, plants, and sites as part of a broader effort to modernize operations across DOE’s laboratory complex.  To advance President Trump’s commitment to Restoring Gold Standard Science, DOE is updating outdated and duplicative operating requirements to give its world-class scientific workforce more time to focus on critical science, energy, and national security missions. These reforms will improve efficiency, strengthen stewardship of taxpayer resources, and help DOE’s National Laboratories, plants, and sites operate with the speed, discipline, and agility their missions demand, while maintaining rigorous safety and security standards.  “America’s National Laboratories are among our nation’s greatest scientific assets and have powered generations of American discovery and innovation,” said U.S. Secretary of Energy Chris Wright. “President Trump has called on DOE to build on that legacy by restoring Gold Standard Science and unleashing the full potential of American ingenuity. By removing unnecessary barriers, we are giving our scientists, engineers, and technicians, more freedom to focus on the critical missions that matter most.” Working with laboratory leaders and subject matter experts, DOE reviewed a targeted set of directives governing day-to-day field operations. Its reforms build on more than three decades of recommendations from Congress, the Government Accountability Office, the National Academies, and other independent reviews that have identified unnecessary complexity in DOE’s directives framework.  DOE is acting on these longstanding recommendations while preserving strong oversight, accountability, and operational excellence—including strong protections for DOE workers, the public, the environment, and the Nation’s nuclear security enterprise.   DOE’s National Laboratories, plants, and sites carry out some of the nation’s most consequential scientific, engineering, and national security missions. Today’s action better aligns their operations with the pace and complexity of today’s missions, giving its scientific workforce more time to develop technologies, strengthen American

Read More »

Energy Secretary Announces Cancellation of Three Proposed National Interest Electric Transmission Corridors

WASHINGTON—U.S. Secretary of Energy Chris Wright today announced that the U.S. Department of Energy (DOE) will not move forward with designating the three proposed National Interest Electric Transmission Corridors (NIETCs) previously selected in December 2024 to advance in the review process. “Extensive review, including public feedback and stakeholder input, made clear that the current designation process for these three proposed transmission corridors should not continue,” said Secretary Wright. “Transmission policy must serve the American people—not special interests or a climate-alarmist agenda that drives up costs, worsens reliability, and disregards the concerns of local communities. The Trump Administration is committed to strengthening America’s electric grid with common-sense policies that prioritize delivering affordable, reliable, and secure electricity to American families and businesses.” The previous administration touted the Lake Erie–Canada Corridor, the Southwestern Grid Connector Corridor, and the Tribal Energy Access Corridor, as a means to advance their Green New Scam agenda and “accelerate decarbonization.” As the process unfolded, the current designation framework proved ineffective in strengthening grid reliability and reducing electricity costs. In some communities, it also contributed to confusion and concern about the scope and intent of NIETC authority. Thanks to President Trump and Secretary Wright, DOE has already taken numerous steps to build new transmission infrastructure and modernize existing infrastructure, including: In October 2025, DOE’s Office of Energy Dominance Financing (EDF) closed a $1.6 billion loan guarantee to AEP Transmission to reconductor and rebuild nearly 5,000 miles of transmission lines across five states.  In February 2026, DOE’s Office of Energy Dominance Financing (EDF) closed $26.5 billion in loans to Southern Company subsidiaries Georgia power and Alabama Power to support generation and grid investments, including more than 1,300 miles of transmission and grid enhancement projects.  In March 2026, DOE’s Office of Electricity (OE) announced the $1.9 billion SPARK funding opportunity to

Read More »

Permian Resources lifts forecast on working interest gains, acquisitions

@import url(‘https://fonts.googleapis.com/css2?family=Inter:wght@100..900&display=swap’); .ebm-page__main h1, .ebm-page__main h2, .ebm-page__main h3, .ebm-page__main h4, .ebm-page__main h5, .ebm-page__main h6 { font-family: Inter; } body { line-height: 150%; letter-spacing: 0.025em; } button, .ebm-button-wrapper { font-family: Inter; } .label-style { text-transform: uppercase; color: var(–color-grey); font-weight: 600; font-size: 0.75rem; } .caption-style { font-size: 0.75rem; opacity: .6; } #onetrust-pc-sdk [id*=btn-handler], #onetrust-pc-sdk [class*=btn-handler] { background-color: #c19a06 !important; border-color: #c19a06 !important; } #onetrust-policy a, #onetrust-pc-sdk a, #ot-pc-content a { color: #c19a06 !important; } #onetrust-consent-sdk #onetrust-pc-sdk .ot-active-menu { border-color: #c19a06 !important; } #onetrust-consent-sdk #onetrust-accept-btn-handler, #onetrust-banner-sdk #onetrust-reject-all-handler, #onetrust-consent-sdk #onetrust-pc-btn-handler.cookie-setting-link { background-color: #c19a06 !important; border-color: #c19a06 !important; } #onetrust-consent-sdk .onetrust-pc-btn-handler { color: #c19a06 !important; border-color: #c19a06 !important; } The leaders of Permian Resources Corp., Midland, have lifted their production and capital spending forecasts for 2026 after recently closing on a $520 million acquisition, exercising an option for a 5,600-acre bolt-on buy and growing its working interest in completed wells more than expected. Permian Resources on July 31 closed on the purchase of about 20,500 acres in the Delaware basin’s Ward County that are largely non-operated and produce about 5,000 boe/d. The land sits adjacent to Permian property but James Walter, co-chief executive officer, told analysts on Aug. 6 that his team have since struck a deal with another operator that will trade some of the acquired bolt-on parcels as well as other acreage with goals of densifying Permian Resources’ holdings and lowering the share of acres that are non-operated or have low working interest. Permian Resources Corp. Permian Resources’ trade in Ward County with another operator is expected to close later this quarter. <!–> ]–> “The trade also increases the number of operating net locations from 50 to 120 while increasing the average lateral length by 20%,” Hickey said. “We view this trade as a true win-win for [Permian Resources] and our counterparty, who

Read More »

Canada rig count down 3 units

The rig count in Canada fell by 3 units to 216 rigs working for the week ended Aug. 7, according to data from Baker Hughes. A 4-rig drop in oil-directed rigs in Canada was partially offset by a 2-unit gain in gas-directed rigs. There were 146 oil-directed rigs working in Canada this week, while those drilling for gas ended the week at 65 units working. The overall US drilling rig count was unchanged this week at 588 rigs working. That number is up 49 units from this time last year. In the US, 3 additional rigs were drilling for oil, bringing the total count to 454. That number is up 43 units from this time last year. The number of gas-directed rigs fell by 3 to 124 working for the week. This time last year, 123 rigs were drilling for gas in the US. There were 572 rigs drilling on US land this week, unchanged from last week and up 48 from the year-ago period. A 1-rig increase in offshore rigs offset a 1-unit decrease in rigs drilling in inland waters. There were 14 rigs drilling offshore and 2 in inland waters this week. Leading the major oil-and gas-producing states was Texas with a 2-unit gain to end the week with 275 rigs working. The count is up 32 units from this time in 2025. Pennsylvania and Wyoming each dropped a rig to bring the respective rig counts to 16 and 15 for the week.

Read More »

Orlen’s Mažeikiai refinery to benefit from renewable electricity

Orlen SA has brought a 42.2-Mw solar photovoltaic (PV) farm online to supply renewable energy that will help to power operations at subsidiary Orlen Lietuva AB’s 10.4-million tonne/year refinery in Mažeikiai, Lithuania. Operable as of Aug. 11 and designed to generate about 45 gigawatt-hours (Gw-hr)/year of electricity, the Mažeikiai solar farm aims to reduce the refinery’s electricity procurement costs by about €4 million/year while supporting Orlen’s goal of increasing the share of renewables across its portfolio, the company said. Located on site across 60 hectares on the refinery’s grounds, the solar farm consists of about 68,000 bifacial photovoltaic modules. Each module is rated at 620 w, the bifacial design of the modules enabling the capture of sunlight on both sides to improve energy output during lower-light conditions on cloudy days, according to Orlen. The solar PV farm’s generation of about 45 Gw-hr of electricity will cover roughly 7% of the Mažeikiai manufacturing complex, where it will be dedicated to supplying power for day-to-day refinery operations, office buildings, and other critical infrastructure at the site. Completed at an overall investment of nearly €35 million, Orlen said the solar farm project received €2.5 million in support from the European Union’s Modernization Fund. Energy transition, efficiency Alongside strengthening the refinery’s energy security by providing an on-site source of reliable electricity, the new solar farm advances Orlen’s commitment to advancing regional energy transition initiatives. “This is an important step towards reducing the environmental impact of our operations and lowering the [Mažeikiai] refinery’s operating costs,” said Dariusz Zonenberg, Orlen Lietuva’s chief executive officer. “The project will increase the share of Orlen Lietuva’s electricity demand met by its own renewable generation, strengthening the company’s competitiveness and supporting the Orlen Group’s long-term strategy,” Zonenberg added. Orlen said the project supports its 2035 strategy to expand renewable energy

Read More »

ADNOC Gas advances its largest-ever gas processing expansion

Abu Dhabi National Oil Co. (ADNOC) subsidiary ADNOC Gas PLC has let a contract to Tecnimont SPA—a subsidiary of Maire SPA—to provide a suite of services for the third phase of the operator’s broader multibillion-dollar, multiphased Rich Gas Development (RGD) project that aims to expand the company’s natural gas processing capacity to meet rising energy demand and secure the United Arab Emirates’ (UAE) reliability as a global energy supplier. As part of the $4.3-billion contract officially revealed on Aug. 10 following intimations to the market in releases dated June 4 and May 20 that withheld the identity of the operator and project, Tecnimont will deliver engineering, procurement, and construction (EPC) services for ADNOC Gas’ RGD Phase 3 expansion involving the addition of a fifth NGL fractionation unit at the Ruwais NGL complex in Abu Dhabi, Maire said. Alongside the NGL fractionation unit designed to separate various hydrocarbon components, as well as treatment and sweetening systems to remove impurities and ensure product quality, Maire confirmed Tecnimont’s scope of work also will cover EPC for a new regeneration gas treatment unit, a propane refrigeration system, ancillary systems, and associated storage installations of the RGD Phase 3 project. Scheduled for completion in 2030, the Phase 3 plant will have an output capacity of 23,000 tonnes/day, equivalent to about 8 million tonnes/year (tpy), according to the service provider. Confirmation of the Phase 3 contract award follows ADNOC Gas’ announcement earlier on Aug. 10 that it had taken final investment decision on both Phase 2 and Phase 3 of the RGD project, including the operator’s separate and concurrent award to Wison Engineering Ltd. for the project’s second phase. As part of the $3.9-billion RGD Phase 3 contract, Wison Engineering will deliver EPC services for a new 670-MMcfd natural gas processing train at the operator’s Habshan

Read More »

It’s final! Judge says HPE’s Juniper acquisition is complete

Originally announced on January 9, 2024, the deal and has undergone public scrutiny ever since, with regulatory reviews in the UK, EU and the US. It was the US that proved to be the final hurdle, with the Justice Department suing to block the deal at first. At the time, the DOJ said reduced competition in the wireless market would be the biggest problem with the proposed buy. In its statement, the agency noted that HPE and Juniper are the second- and third-largest providers, respectively, of enterprise-grade WLAN solutions in the U.S. behind market leader Cisco. But those issues were ultimately settled in June 2025, and HPE has gone on to integrate Juniper’s networking technology. Most recently, it announced a raft of new products, including HPE Juniper Networking QFX switches aimed at inferencing and scale-up architecture. It also deepened integration of its Juniper Networking data center switching and operations into its Mist AI engine and launched a unified, AI-native SASE platform.

Read More »

Google, Microsoft and Nvidia back 800V DC standard for AI data centers

The savings over alternating current (AC), the current power standard, are considerable. With AC, there are 4 wires while DC has two. So there is considerable wiring savings in an all-DC facility. Also, with higher voltage comes a lower current and current is what generates heat. So data centers that can run on 800VDC can run cooler. That translates to a 50% to 80% reduction in copper usage and an 8% to 12% reduction in annual energy-related OpEx through lower conversion and distribution losses. AI-first facilities can see a $4 million to $8 million in CapEx savings per 10 MW build by reducing upstream AC. For a one-gigawatt data center, you’re saving a several million pounds of copper wire. The push reflects a fundamental change in data-center power requirements. AI accelerators are being deployed in increasingly dense configurations, driving power consumption per rack higher and making traditional low-voltage AC distribution more difficult to scale.

Read More »

Five takeaways from Cisco’s Q4 and what they mean for IT pros

1. Agentic AI is driving a “networking supercycle” During the call, CEO Chuck Robbins repeatedly emphasized that accelerating agentic AI adoption is fueling a long-term “networking supercycle.” For years, network traffic was predictable: client-to-server or standard east-west data center traffic. Agentic AI upends those legacy traffic models. Autonomous AI agents interact continuously with application programming interfaces (API), databases, vector search engines, and other agents, driving massive increases in lateral bandwidth requirements and imposing strict low-latency constraints. Furthermore, as AI models grow in size, physical data center boundaries are proving insufficient. Hyperscalers and large enterprises are adopting scale-across architectures that link multiple physical data centers, enabling distributed GPUs to operate as a single logical cluster. Cisco noted that network traffic in scale-across environments is roughly 14 times higher than in traditional data center interconnects. What it means for IT pros: If your team still treats network capacity planning as an annual incremental upgrade, you will be left behind. Agentic workflows will overwhelm LANs, WANs, and data center networks with unprecedented volumes of multidirectional traffic. Network architects must immediately evaluate non-blocking topologies, high-density 400G/800G switching, and deterministic networking to prevent enterprise AI initiatives from stalling at the transport layer.

Read More »

North American data center vacancy holds at just 1%

The shift is being driven largely by the industry’s growing need for electricity, land and infrastructure. Traditional markets such as Northern Virginia have become increasingly difficult to expand because of power constraints, land availability and lengthy utility interconnection timelines. But JLL cautioned that the industry’s continued expansion will depend increasingly on winning public support, at a time of significant pushback from residents. “Supporting the next phase of growth will depend on building trust, addressing local concerns and delivering lasting benefits to host communities,” the report said. In addition to residential pushback, power availability is becoming increasingly scarce. In primary data center markets, the average wait for a grid connection can exceed four years, pushing operators toward on-site generation, battery storage and other alternatives.

Read More »

IT infrastructure shortages are real and lasting. Here’s how to cope

Look at alternatives, including AMD and cloud solutions, while staying mindful of how it all plays together. You may not be able to get Nvidia GPUs, but AWS, Azure, and Oracle Cloud have them, Kimball notes. Be strategic, perhaps by using cloud offerings to handle certain tuning or inference workloads, then bringing them back in-house when appropriate. “Have a better understanding of what absolutely has to be on prem and what can be in the cloud,” he says. That’s good advice, says Backblaze’s Thomas. When it comes to AI, think about performance tiers and the range of use cases you have. They don’t all need top-tier performance. “People get wrapped around axle of needing the top end. There’s a lot of flexibility in the edges, innovation in different hardware and software,” Thomas says. Gartner likewise advises companies to increase configuration flexibility and expand sourcing paths. That may include buying from secondary markets and lease-return programs to preserve continuity with existing infrastructure until the shortages pass, Forest says. Get started somewhere Even if you can’t acquire or have to wait for the infrastructure you need, don’t let that keep you from getting started with AI or other modernization projects. Options include public cloud and neocloud providers, Anderson says. WWT also provides capacity in its own lab so customers can get started with proof-of-concept projects. “Don’t just throw your hands up. We can help you find access to capacity,” Anderson says. “Production-scale AI may be delayed, but don’t let that derail your strategy.” Colocation providers may likewise be an option, especially if enterprises are struggling to acquire high-end networking equipment. Networking is a key value proposition for colocation providers, in that they have built-in connections to various cloud providers and other ecosystem players. Equinix, for example, has 280 data centers in 77 metropolitan

Read More »

Polish data center plans to send its waste heat to the neighbors

As Europe swelters in a heatwave, residents probably don’t want to hear about ways to make their homes even hotter, but that’s what Polish property developer Citylink is talking about, with plans to dump waste heat from a new data center in Wrocław into the municipal district heating network. Citylink is designing the data center so that heat from servers can be recovered instead of being dissipated via cooling systems — and as the data center grows, any increase in computing power will mean more energy available for recovery. The collaboration with local power company Kogeneracja will provide “valuable experience in designing and operating modern data centers, with a particular focus on infrastructure dedicated to AI nodes,” said Michał Starybrat, development director at Citylink.

Read More »

Microsoft will invest $80B in AI data centers in fiscal 2025

And Microsoft isn’t the only one that is ramping up its investments into AI-enabled data centers. Rival cloud service providers are all investing in either upgrading or opening new data centers to capture a larger chunk of business from developers and users of large language models (LLMs).  In a report published in October 2024, Bloomberg Intelligence estimated that demand for generative AI would push Microsoft, AWS, Google, Oracle, Meta, and Apple would between them devote $200 billion to capex in 2025, up from $110 billion in 2023. Microsoft is one of the biggest spenders, followed closely by Google and AWS, Bloomberg Intelligence said. Its estimate of Microsoft’s capital spending on AI, at $62.4 billion for calendar 2025, is lower than Smith’s claim that the company will invest $80 billion in the fiscal year to June 30, 2025. Both figures, though, are way higher than Microsoft’s 2020 capital expenditure of “just” $17.6 billion. The majority of the increased spending is tied to cloud services and the expansion of AI infrastructure needed to provide compute capacity for OpenAI workloads. Separately, last October Amazon CEO Andy Jassy said his company planned total capex spend of $75 billion in 2024 and even more in 2025, with much of it going to AWS, its cloud computing division.

Read More »

John Deere unveils more autonomous farm machines to address skill labor shortage

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Self-driving tractors might be the path to self-driving cars. John Deere has revealed a new line of autonomous machines and tech across agriculture, construction and commercial landscaping. The Moline, Illinois-based John Deere has been in business for 187 years, yet it’s been a regular as a non-tech company showing off technology at the big tech trade show in Las Vegas and is back at CES 2025 with more autonomous tractors and other vehicles. This is not something we usually cover, but John Deere has a lot of data that is interesting in the big picture of tech. The message from the company is that there aren’t enough skilled farm laborers to do the work that its customers need. It’s been a challenge for most of the last two decades, said Jahmy Hindman, CTO at John Deere, in a briefing. Much of the tech will come this fall and after that. He noted that the average farmer in the U.S. is over 58 and works 12 to 18 hours a day to grow food for us. And he said the American Farm Bureau Federation estimates there are roughly 2.4 million farm jobs that need to be filled annually; and the agricultural work force continues to shrink. (This is my hint to the anti-immigration crowd). John Deere’s autonomous 9RX Tractor. Farmers can oversee it using an app. While each of these industries experiences their own set of challenges, a commonality across all is skilled labor availability. In construction, about 80% percent of contractors struggle to find skilled labor. And in commercial landscaping, 86% of landscaping business owners can’t find labor to fill open positions, he said. “They have to figure out how to do

Read More »

2025 playbook for enterprise AI success, from agents to evals

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More 2025 is poised to be a pivotal year for enterprise AI. The past year has seen rapid innovation, and this year will see the same. This has made it more critical than ever to revisit your AI strategy to stay competitive and create value for your customers. From scaling AI agents to optimizing costs, here are the five critical areas enterprises should prioritize for their AI strategy this year. 1. Agents: the next generation of automation AI agents are no longer theoretical. In 2025, they’re indispensable tools for enterprises looking to streamline operations and enhance customer interactions. Unlike traditional software, agents powered by large language models (LLMs) can make nuanced decisions, navigate complex multi-step tasks, and integrate seamlessly with tools and APIs. At the start of 2024, agents were not ready for prime time, making frustrating mistakes like hallucinating URLs. They started getting better as frontier large language models themselves improved. “Let me put it this way,” said Sam Witteveen, cofounder of Red Dragon, a company that develops agents for companies, and that recently reviewed the 48 agents it built last year. “Interestingly, the ones that we built at the start of the year, a lot of those worked way better at the end of the year just because the models got better.” Witteveen shared this in the video podcast we filmed to discuss these five big trends in detail. Models are getting better and hallucinating less, and they’re also being trained to do agentic tasks. Another feature that the model providers are researching is a way to use the LLM as a judge, and as models get cheaper (something we’ll cover below), companies can use three or more models to

Read More »

OpenAI’s red teaming innovations define new essentials for security leaders in the AI era

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More OpenAI has taken a more aggressive approach to red teaming than its AI competitors, demonstrating its security teams’ advanced capabilities in two areas: multi-step reinforcement and external red teaming. OpenAI recently released two papers that set a new competitive standard for improving the quality, reliability and safety of AI models in these two techniques and more. The first paper, “OpenAI’s Approach to External Red Teaming for AI Models and Systems,” reports that specialized teams outside the company have proven effective in uncovering vulnerabilities that might otherwise have made it into a released model because in-house testing techniques may have missed them. In the second paper, “Diverse and Effective Red Teaming with Auto-Generated Rewards and Multi-Step Reinforcement Learning,” OpenAI introduces an automated framework that relies on iterative reinforcement learning to generate a broad spectrum of novel, wide-ranging attacks. Going all-in on red teaming pays practical, competitive dividends It’s encouraging to see competitive intensity in red teaming growing among AI companies. When Anthropic released its AI red team guidelines in June of last year, it joined AI providers including Google, Microsoft, Nvidia, OpenAI, and even the U.S.’s National Institute of Standards and Technology (NIST), which all had released red teaming frameworks. Investing heavily in red teaming yields tangible benefits for security leaders in any organization. OpenAI’s paper on external red teaming provides a detailed analysis of how the company strives to create specialized external teams that include cybersecurity and subject matter experts. The goal is to see if knowledgeable external teams can defeat models’ security perimeters and find gaps in their security, biases and controls that prompt-based testing couldn’t find. What makes OpenAI’s recent papers noteworthy is how well they define using human-in-the-middle

Read More »