Stay Ahead, Stay ONMINE

On-device translation: Why smaller, specialized AI models win

The obsession with “bigger is better” AI has led many consumers to assume that more parameters mean better performance. This isn’t always true, particularly for specific use cases. Smaller AI models that can run on the device itself, rather than in the cloud, are a better fit for use cases like translation apps. While your […]

The obsession with “bigger is better” AI has led many consumers to assume that more parameters mean better performance. This isn’t always true, particularly for specific use cases.

Smaller AI models that can run on the device itself, rather than in the cloud, are a better fit for use cases like translation apps.

While your laptop may handle the heavy computational requirements of multi-purpose AI models trained on models with billions of parameters, your mobile phone and smartwatch may struggle.

In the case of translation, and in more constrained use cases, some developers are building smaller, purpose-built translation models that do one thing well.

Tether is one of the companies building in this direction, with an AI team that has developed dedicated, resource-optimized translation models designed to run entirely on-device.

Your translations should stay on your device

A translation app directly gains access to your conversation. Just like your mobile chat app, it handles information, some of which may be confidential. This is the same even when the translator doesn’t communicate directly with your chat application, for example, by copying and pasting text into the translator.

However, most translation apps are connected to the cloud. That means your text, whether it’s a business contract, a medical record, or a private conversation, leaves your device every time you use it. Local translation eliminates that.

Consequently, developers are exploring online translators and embracing a local-first, offline approach as the only viable way to create effective translators without a single point of failure.

When less is better

Developing models dedicated to a single purpose allows developers to trim down the size and execution costs for the models. Prototypes designed this way are usually only a few megabytes (MB) in size and can run on regular devices.

As demonstrated by Tether’s AI team, Tether’s Bergamot-compatible models require only 21-35 MB per language pair and efficiently translate inputs at ~46ms per sentence, which is ~78 times faster than the 2-billion-parameter Salamandra model.

The modularity of these dedicated models is their biggest appeal. They are lightweight, edge-optimized, and flexible enough to fit into heterogeneous systems with a minimal footprint. This means easier integrations and more practicality.

These models are ideal for local AI due to their low compute requirements. They can be installed and run on any device, including mid-range mobile phones and IoT devices. On-premises integration for this use case offers even greater advantages for users.

Unifying efficient, local-first translation models

As considerable progress has already been made in developing local and offline translation models, Tether is leading the next stage: Application. Tether’s QVAC SDK unifies these models into a single wrapper and provides a directory of models that allows users and developers to choose a preferred model for each language pair.

Tether’s QVAC SDK simplifies Neural Machine Translation (NMT) model implementation for developers and regular users by providing prebuilt modules that enable anyone to select, deploy, and manage intelligent language translators in their applications.

It packages language pairs as dependencies that can be loaded via simple import statements and used in code to handle language translation requests.

The SDK provides primitives for intelligent language translation, supports practical scenarios (single sentences and multiple sentences via batch translation), and a fallback for the unlikely situation where the specialized, lightweight NMT models are unable to meet a developer’s or users’ needs.

The fallback is an LLM-based translation framework for training new language models or running direct translations.

Scaling efficiently to hundreds of languages

The translation component in the QVAC SDK supports multiple language pairs, reducing the number of packages required to run a translation system across hundreds of languages. For instance, a developer would normally need 2 language pairs to create a complete English-to-Chinese translation system (ENG-ZH and ZH-ENG).

To support 26 languages, this dramatically grows to 650 translation directions. The English-pivot model in QVAC SDK reduces this to just 50 language pairs for a 26-language translator.

Tether’s QVAC SDK brings resource-efficient, local, and private translation to everyone’s doorstep. It complements the progress made with the NMT technology by scaling adoption for limitless application scenarios.

Tether’s vision to support local and edge-first AI development also extends to its other technology offerings, including Brain-Computer Interfaces (BCI). Brain OS, the Brain Operating System designed by Tether AI Research engineering team and built on top of Tether’s QVAC AI platform, aims to create an open-source brain operating system that connects to the user’s personal BCI. The idea is that our most important data (our thoughts) should always remain private and owned by us.

Start building intelligent local and edge-first applications with multi-language support. Check out the QVAC repo. 

Shape
Shape
Stay Ahead

Explore More Insights

Stay ahead with more perspectives on cutting-edge power, infrastructure, energy,  bitcoin and AI solutions. Explore these articles to uncover strategies and insights shaping the future of industries.

Shape

Intel unwraps three-pronged architecture strategy to go after agentic AI

While Diamond Rapids handles the general-purpose side of the equation, Crescent Island is aimed directly at inference. Intel describes the accelerator as a relatively low-power, air-cooled GPU designed to deliver higher token throughput while accommodating larger models, longer context windows and more concurrent AI agents. Crescent Island uses 32 Xe3P-based

Read More »

DOE Selects Community Partners to Receive Waste to Energy and Materials Recovery Technical Assistance

WASHINGTON—The U.S. Department of Energy’s Alternative Fuels and Feedstocks Office (AFFO) and the National Laboratory of the Rockies (NLR) have selected recipients for the FY26 Waste to Energy and Materials Technical Assistance program. Through this program, NLR will provide free guidance to state, local, and Tribal governments to use new technologies that turn waste into energy or recover valuable materials like critical minerals.  The program aims to help local officials create sensible solutions for their waste management issues, fill knowledge gaps, and plan and carry out implementation approaches that fit their communities. This year, the program has expanded to include additional municipal solid waste streams like electronics, industrial wastewater, and other byproducts.  Now in its sixth year, the technical assistance program has supported 67 entities in 31 states and territories. FY26 selections include: Community Name American Samoa Power Authority City of Boise, Idaho Cherokee Nation Natural Resources, Oklahoma Village of Coal Valley, Illinois Guam Energy Office Hudson Valley Regional Council, New York Kodiak Island Borough, Alaska Los Angeles County Public Works, California Metlakatla Indian Community, Arkansas Township of Montclair, New Jersey City of New Bedford, Massachusetts Oregon Department of Energy, Oregon South Central Regional Council of Governments, Connecticut Thompson Township, Pennsylvania Ulster County Resource Recovery Agency, New York Washington State Department of Commerce, Office of Renewable Fuels To learn more about the technical assistance program, visit NLR’s Waste to Energy and Materials Technical Assistance for State, Local, and Tribal Governments webpage. If you have questions, please see frequently asked questions or contact the Waste to Energy and Materials Technical Assistance Team.

Read More »

Santos targets Q4 2026 FID for Papua LNG plant

Santos Ltd. is on track to take fourth-quarter 2026 final investment decision (FID) on the 5.6 million tonne/year (tpy) Papua LNG plant at Caution Bay, Papua New Guinea, with project financing and government-led development discussions advancing. At plateau, Papua LNG would contribute about 1 million tpy of Santos equity LNG and roughly 11 million boe/year of equity oil, the company said in its first-half 2026 earnings report and call. Papua LNG would use 4 million tpy of production from new electric liquefaction trains and as much as 2 million tpy of tolling production from ExxonMobil Corp.’s already operating 8-million tpy PNG LNG plant, in which Santos is also a partner. Santos recently took FID on its PNG LNG oil infill drilling campaign, and expects to start drilling fourth-quarter 2026. Santos said it has several options to backfill PNG LNG production if Papua LNG does not proceed but emphasized that all parties remain focused on reaching a Papua LNG FID this year. The company cited Muruk, P’nyang, and Usano as possible resources for such backfill. Muruk has estimated natural gas resources of 1-3 tcf and P’nyang estimated recoverable reserves of 4.36 tcf. Usano, in the PD-L2 production license area, is primarily an oil project, with an estimated 85 million bbl of oil in place but would produce associated gas as well. Santos plans to drill a test well on it in early 2028. TotalEnergies SE holds 40.1% interest in Papua LNG and it the project’s operator. ExxonMobil holds 37.1% interest, with Santos and the state holding the bulk of the balance. Santos equity is 17.7-22.8% depending on government exercise of its back-in rights. ENEOS (formerly JX Nippon) holds a minor participating interest.

Read More »

North American rig count drops 8 units, erasing last week’s gain

The rig count in North America is down 8 units this week, according to data from Baker Huges Inc. With 804 rigs running across North America for the week ended Aug. 21, the drop erased the previous week’s 8-unit gain. There were 5 fewer rigs drilling in the US this week for a total of 588. The count is 50 more than were drilling during the same period last year. A 2-unit drop in offshore rigs left 10 working this week. One fewer rig was drilling in inland waters, leaving 2 still working. The number of rigs drilling on land decreased by 2 to 576. That count is up 53 from the same period in 2025. Three fewer rigs were oil-directed in the US and its waters this week for a total of 452. There were 127 gas-directed rigs working, down one from last week. The number of unclassified rigs working this week decreased by 1 unit to 9. Of the major US oil and gas producing states, Texas saw the largest increase. Four rigs were added to the state’s total this week to bring the count to 281, 41 more than were drilling during the same period last year. New Mexico and Louisiana each dropped 3 rigs to end the week with counts of 96 and 35, respectively. Wyoming’s rig count fell by a single unit this week to leave 9 rigs working. The overall rig count in Canada fell by 3 units to 216. The count is up 36 units from this time a year ago. Of those 216 rigs working, 148 were drilling for oil, down 3 from last week. The number of gas-directed rigs in Canada was unchanged at 65. Three units were unclassified, unchanged from last week.

Read More »

Ring’s 2027 target: 10% growth for 10% less

Boosted by an increase in horizontal drilling across its Central Basin Platform (CBP) operations, the leaders of Ring Energy Inc., The Woodlands, Tex., expect a big pop in the company’s 2027 financials. Speaking Aug. 18 at the EnerCom Denver conference, chairman and chief executive officer Paul McKinney said the Permian basin-focused operator has “an incredible runway of high-return opportunities” in the CBP using technologies refined by operators in the Midland and Delaware basins on either side of Ring’s holdings. Recent developments, he said, have made it easier for Ring and others active in the CBP, which has shallower reservoirs, to drill longer wells. Two years ago, half of the wells Ring drilled were horizontal. This year, that figure is on pace to be 81%. The length of new wells is similarly shifting to being at least 1.5 miles: In 2024, new wells of that length accounted for only 5% of Ring’s activity but that will be 70% this year. Those advancements are set to create a big payoff for Ring, which had total production of just under 20,000 boe/d in the second quarter. “The capital is kind of the story,” McKinney told EnerCom attendees. “We believe that we will deliver 10% production growth for 10% less capital in 2027 […] All this means meaningful upside in adjusted free cash flow. It means a significant increase in earnings.”

Read More »

Federal judge allows Sable Offshore to continue California pipeline operations

Despite affirming the jurisdictional shift, Wilson also ordered Sable to pay $1.5 million for violating a federal consent decree. Through its acquisition of the assets, Sable assumed obligations under the decree, including management and reporting requirements and provisions requiring state waivers before restarting operations. “Sable has violated the express provisions of the consent decree, without justification,” Wilson wrote. The judge said California’s proposed injunction “is not the proper remedy.” For one, he said, “the consent decree has been modified to replace OSFM as the regulatory authority with PHMSA, and the pre-restart requirements of the State Waivers are no longer applicable. Nor, too, are OSFM’s approval of a Restart Plan or authorization. PHMSA, the current regulator, has authorized Sable to restart the pipeline. Therefore, Sable is no longer in violation of the Consent Decree, and proactive, injunctive relief is inappropriate,” Wilson wrote. “Rather, the appropriate penalty for Sable’s violations is dictated by the consent decree.” Sable resumed transporting crude oil from the Santa Ynez Unit (SYU) through SYPS in March under the DPA order. The order and company statements indicate gross oil throughput is expected to reach about 50,000 b/d following ramp-up. Current production from six wells is estimated at about 6,000 b/d. SYPS has capacity of up to 200,000 b/d.

Read More »

US threatens sanctions against countries, companies buying Iranian oil

@import url(‘https://fonts.googleapis.com/css2?family=Inter:wght@100..900&display=swap’); .ebm-page__main h1, .ebm-page__main h2, .ebm-page__main h3, .ebm-page__main h4, .ebm-page__main h5, .ebm-page__main h6 { font-family: Inter; } body { line-height: 150%; letter-spacing: 0.025em; } button, .ebm-button-wrapper { font-family: Inter; } .label-style { text-transform: uppercase; color: var(–color-grey); font-weight: 600; font-size: 0.75rem; } .caption-style { font-size: 0.75rem; opacity: .6; } #onetrust-pc-sdk [id*=btn-handler], #onetrust-pc-sdk [class*=btn-handler] { background-color: #c19a06 !important; border-color: #c19a06 !important; } #onetrust-policy a, #onetrust-pc-sdk a, #ot-pc-content a { color: #c19a06 !important; } #onetrust-consent-sdk #onetrust-pc-sdk .ot-active-menu { border-color: #c19a06 !important; } #onetrust-consent-sdk #onetrust-accept-btn-handler, #onetrust-banner-sdk #onetrust-reject-all-handler, #onetrust-consent-sdk #onetrust-pc-btn-handler.cookie-setting-link { background-color: #c19a06 !important; border-color: #c19a06 !important; } #onetrust-consent-sdk .onetrust-pc-btn-handler { color: #c19a06 !important; border-color: #c19a06 !important; } US Treasury Secretary Scott Bessent Aug. 24 threatened secondary sanctions on any nation or entity maintaining economic ties with Iran, including buying its oil. He specifically warned Tehran’s trading partners (China, Turkey, and the UAE) to sever economic links or face direct US financial retaliation.  The Treasury plan, Operation Economic Outcast, does not immediately impose the secondary sanctions; it was framed as a deadline-driven warning to give violators time to comply. Bessent said President Trump was already contacting world leaders to request formal cooperation in isolating Tehran in hopes of ending the 6-month-plus old war. The sanctions package continues to crack down the ‘shadow fleet’ by expanding tracking and blacklisting the oil tankers, shipping insurers, and others that help smuggle Iranian crude. It also targets financial intermediaries that help channel oil sales into usable revenues. China, the largest buyer of Iranian crude oil, issued a formal order directing Chinese companies to disregard US sanctions. China purchases 80-90% of Iran’s exported oil. Beijing also vowed to take all “necessary measures” to defend its energy security and trading rights.

Read More »

IBM unveils dual-architecture processor to run Arm-native apps on Z mainframes

“These caches have enormously low latency, and that is one of the key reasons and key engineering choices to support the performance and scalability of enterprise workloads, very data-intensive workloads like databases and transactions,” Jacobi said. “In addition, we have an on-chip data processing unit for IO acceleration and dedicated AI accelerators as well as accelerators for data compression, cryptography and data sorting.” One of the biggest takeaways from this processor announcement is that the enormous catalog of software already built for Arm becomes accessible on a mainframe without anyone having to port it first, notes Matt Kimball, senior datacenter analyst at Moor Insights & Strategy, in a research note about the news. Still, “this is a 2027 conversation, and with no date, supported software list, or Arm licensing treatment, the work now is inventory and scenario planning rather than financial modeling,” Kimball wrote.

Read More »

PJM’s New Data Center Power Equation

PJM Interconnection has now filed one of the most consequential proposed changes yet in the relationship between data centers and the electric grid. Rather than simply treating a new hyperscale or AI facility like any other customer whose demand will be backed through regional capacity procurement, PJM is proposing a framework under which the largest new loads would need to be supported by new capacity, have their needs covered through the Reliability Backstop Procurement, or face potential curtailment when the regional power system is short of supply. The approach has been developing since PJM launched its Critical Issue Fast Path process for large loads in 2025, but it became substantially more concrete in late July and August 2026. PJM filed its proposed Reliability Backstop Procurement with FERC on July 31 and began accepting applications that day for its FERC-approved Expedited Interconnection Track. On Aug. 13, PJM filed its proposed Interim Resource Adequacy Service, or IRAS, along with the Large Load Registry that would support it. The immediate numbers explain the urgency. PJM’s July 2026 capacity auction for the 2028/2029 delivery year procured 138,318 MW of unforced capacity through the centralized auction. Even after including Fixed Resource Requirement resources, however, PJM came up 6,831 MW short of its reliability requirement. The auction cleared at the FERC-approved $325/MW-day price cap. It was the second consecutive auction in which the PJM region failed to procure its full reliability requirement, something that had not happened before these two auctions. That gap is occurring while demand continues to accelerate. PJM’s 2026 long-term forecast projects summer peak demand growing at an average 3.6% annually over the next decade, compared with just 0.3% in the comparable forecast issued in 2021. Summer peak demand is projected to rise by nearly 66 GW over 10 years. Data centers are

Read More »

Zayo, NVIDIA Build the Long-Haul Backbone for Distributed AI

The data center industry’s increasingly power-first approach to site selection has created a follow-on question: Once the megawatts are found, is there enough network infrastructure to make the site useful at AI scale? Zayo and NVIDIA are putting real infrastructure behind that question. Zayo said it is working with NVIDIA to expand network capacity supporting AI factories across North America, including an 8,000-route-mile program targeting some of the fastest-growing AI corridors in the United States. The project encompasses six new long-haul routes along with overbuilds of existing network across 10 high-demand corridors. The announcement arrives as AI data center development moves beyond the largest established hubs toward markets where power and land may be more readily available, but fiber capacity cannot necessarily be taken for granted. That geography is increasingly important. NVIDIA has separately developed “scale-across” networking technology designed to allow AI infrastructure distributed among different buildings — or even data centers separated by hundreds of kilometers — to operate as a more unified computing environment. Put together, the developments suggest that networking is becoming inseparable from the AI factory buildout itself. Power may determine where the next generation of AI infrastructure can be built. Fiber will increasingly determine how effectively those sites can participate in the larger AI ecosystem. Fiber Follows the Power Zayo CEO Steve Smith said AI demand is changing both where network infrastructure is needed and how aggressively capacity must be deployed ahead of development. “AI is fundamentally reshaping where and how network infrastructure needs to be built across the U.S.,” Smith said. The company’s 8,000-mile program is more nuanced than that top-line number might suggest. Zayo disclosed in April that the expansion includes approximately 3,000 route miles across six new long-haul routes, plus more than 5,000 route miles of overbuilds across 10 existing corridors. Zayo

Read More »

Southern’s 17 GW Pipeline Puts AI Power Demand Into Utility Math

The headline number from Southern Company’s latest earnings report is hard to miss: electricity use by data centers across the utility’s system increased 55% in the second quarter compared with a year earlier. But the more consequential numbers may be the ones sitting behind it. Southern now has more than 1.2 GW of operating data center load, up by more than 500 MW from a year ago. At the same time, its electric utilities have signed contracts and large-load agreements totaling more than 17 GW by the mid-2030s, with another 8 GW in late-stage development and a prospective pipeline of large industrial and data center projects exceeding 75 GW. That leaves an enormous gap between the data center megawatts consuming electricity today and the load Southern has contractually positioned itself to serve during the next decade. For the data center industry, that gap may be the most important part of Southern’s second-quarter story. It offers a look at how utilities are beginning to convert the AI infrastructure boom from forecasts and campus announcements into contracts, generation procurement, transmission investment and eventually energized capacity. From Contracts to Megawatts Southern added roughly 6 GW of contracted large load during the quarter alone. Alabama Power signed three projects representing about 3 GW, while Georgia Power reached a 25-year agreement to serve OpenAI’s planned project in Effingham County near Savannah. That facility is expected to require approximately 3.2 GW and begin taking electric service in phases in 2028. The numbers nevertheless require an important distinction. Seventeen gigawatts contracted does not mean 17 GW will suddenly appear on Southern’s grid. Large data center campuses ramp gradually, often over several years, and Southern executives acknowledged that actual customer ramp schedules do not always match the assumptions made when projects are first approved. CEO Chris Womack said

Read More »

PORTS-Pike Takes Shape as an 8-GW AI Infrastructure Model

Back on March 31, 2026, we discussed we discussed SoftBank and SB Energy’s plans to redevelop the former Portsmouth Gaseous Diffusion Plant site near Piketon as a 10-GW artificial intelligence data center campus supported by almost an equal amount of new power generation. At the time, the plan called for as much as 10 GW of new generation, including 9.2 GW of natural gas capacity, along with approximately $4.2 billion of high-voltage transmission infrastructure developed with AEP Ohio. An initial 800-MW data center phase was targeted for service in 2028. The March story was notable because Pike County appeared to offer a preview of a new model for building hyperscale infrastructure: develop the generation, transmission and data center simultaneously rather than wait for an increasingly congested regional grid to deliver multiple gigawatts of capacity. Not to mention the reuse of a brownfield site with the encouragement of the federal government. Since then, almost every important part of the project has moved forward, and on August 17, the most consequential missing pieces fell into place. NVIDIA announced that it will become the exclusive AI compute infrastructure provider for the PORTS-Pike Technology Campus. OpenAI will be the data center customer, signing a 20-year lease with SB Energy for approximately 8 GW of IT capacity. NVIDIA will invest another $1.5 billion in SB Energy and provide credit support for the land, power and shell infrastructure behind an initial 4.25 GW of IT load, with an option covering approximately another 3.75 GW. The Securities and Exchange Commission filing accompanying the announcement makes the financial commitment even more significant. NVIDIA disclosed that its aggregate payment obligation associated with its initial commitment is capped at $105 billion. That is not a conventional capital commitment to spend $105 billion building the campus, nor is it simply a

Read More »

Nvidia scales back financing guarantee for OpenAI data center

Nvidia is scaling back a proposed financial guarantee tied to a massive OpenAI data center project in Ohio, reducing its initial commitment from as much as $250 billion to less than $120 billion, according to report in the Wall Street Journal. Earlier this month, Nvidia announced partnerships with major financial firms including Apollo Global Management, BlackRock, Blackstone, Brookfield Asset Management, Goldman Sachs and KKR, aimed at mobilizing more than $500 billion in capital for AI computing infrastructure. The change represents a significant restructuring of Nvidia’s role in financing the planned facility, which is being developed by SB Energy, a subsidiary of SoftBank. Under the revised arrangement, Nvidia would guarantee financing for the project’s first phase, representing roughly 5 gigawatts of capacity, or half of the total proposed capacity. Financing for the remaining capacity would be considered separately at a later stage.

Read More »

Microsoft will invest $80B in AI data centers in fiscal 2025

And Microsoft isn’t the only one that is ramping up its investments into AI-enabled data centers. Rival cloud service providers are all investing in either upgrading or opening new data centers to capture a larger chunk of business from developers and users of large language models (LLMs).  In a report published in October 2024, Bloomberg Intelligence estimated that demand for generative AI would push Microsoft, AWS, Google, Oracle, Meta, and Apple would between them devote $200 billion to capex in 2025, up from $110 billion in 2023. Microsoft is one of the biggest spenders, followed closely by Google and AWS, Bloomberg Intelligence said. Its estimate of Microsoft’s capital spending on AI, at $62.4 billion for calendar 2025, is lower than Smith’s claim that the company will invest $80 billion in the fiscal year to June 30, 2025. Both figures, though, are way higher than Microsoft’s 2020 capital expenditure of “just” $17.6 billion. The majority of the increased spending is tied to cloud services and the expansion of AI infrastructure needed to provide compute capacity for OpenAI workloads. Separately, last October Amazon CEO Andy Jassy said his company planned total capex spend of $75 billion in 2024 and even more in 2025, with much of it going to AWS, its cloud computing division.

Read More »

John Deere unveils more autonomous farm machines to address skill labor shortage

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Self-driving tractors might be the path to self-driving cars. John Deere has revealed a new line of autonomous machines and tech across agriculture, construction and commercial landscaping. The Moline, Illinois-based John Deere has been in business for 187 years, yet it’s been a regular as a non-tech company showing off technology at the big tech trade show in Las Vegas and is back at CES 2025 with more autonomous tractors and other vehicles. This is not something we usually cover, but John Deere has a lot of data that is interesting in the big picture of tech. The message from the company is that there aren’t enough skilled farm laborers to do the work that its customers need. It’s been a challenge for most of the last two decades, said Jahmy Hindman, CTO at John Deere, in a briefing. Much of the tech will come this fall and after that. He noted that the average farmer in the U.S. is over 58 and works 12 to 18 hours a day to grow food for us. And he said the American Farm Bureau Federation estimates there are roughly 2.4 million farm jobs that need to be filled annually; and the agricultural work force continues to shrink. (This is my hint to the anti-immigration crowd). John Deere’s autonomous 9RX Tractor. Farmers can oversee it using an app. While each of these industries experiences their own set of challenges, a commonality across all is skilled labor availability. In construction, about 80% percent of contractors struggle to find skilled labor. And in commercial landscaping, 86% of landscaping business owners can’t find labor to fill open positions, he said. “They have to figure out how to do

Read More »

2025 playbook for enterprise AI success, from agents to evals

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More 2025 is poised to be a pivotal year for enterprise AI. The past year has seen rapid innovation, and this year will see the same. This has made it more critical than ever to revisit your AI strategy to stay competitive and create value for your customers. From scaling AI agents to optimizing costs, here are the five critical areas enterprises should prioritize for their AI strategy this year. 1. Agents: the next generation of automation AI agents are no longer theoretical. In 2025, they’re indispensable tools for enterprises looking to streamline operations and enhance customer interactions. Unlike traditional software, agents powered by large language models (LLMs) can make nuanced decisions, navigate complex multi-step tasks, and integrate seamlessly with tools and APIs. At the start of 2024, agents were not ready for prime time, making frustrating mistakes like hallucinating URLs. They started getting better as frontier large language models themselves improved. “Let me put it this way,” said Sam Witteveen, cofounder of Red Dragon, a company that develops agents for companies, and that recently reviewed the 48 agents it built last year. “Interestingly, the ones that we built at the start of the year, a lot of those worked way better at the end of the year just because the models got better.” Witteveen shared this in the video podcast we filmed to discuss these five big trends in detail. Models are getting better and hallucinating less, and they’re also being trained to do agentic tasks. Another feature that the model providers are researching is a way to use the LLM as a judge, and as models get cheaper (something we’ll cover below), companies can use three or more models to

Read More »

OpenAI’s red teaming innovations define new essentials for security leaders in the AI era

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More OpenAI has taken a more aggressive approach to red teaming than its AI competitors, demonstrating its security teams’ advanced capabilities in two areas: multi-step reinforcement and external red teaming. OpenAI recently released two papers that set a new competitive standard for improving the quality, reliability and safety of AI models in these two techniques and more. The first paper, “OpenAI’s Approach to External Red Teaming for AI Models and Systems,” reports that specialized teams outside the company have proven effective in uncovering vulnerabilities that might otherwise have made it into a released model because in-house testing techniques may have missed them. In the second paper, “Diverse and Effective Red Teaming with Auto-Generated Rewards and Multi-Step Reinforcement Learning,” OpenAI introduces an automated framework that relies on iterative reinforcement learning to generate a broad spectrum of novel, wide-ranging attacks. Going all-in on red teaming pays practical, competitive dividends It’s encouraging to see competitive intensity in red teaming growing among AI companies. When Anthropic released its AI red team guidelines in June of last year, it joined AI providers including Google, Microsoft, Nvidia, OpenAI, and even the U.S.’s National Institute of Standards and Technology (NIST), which all had released red teaming frameworks. Investing heavily in red teaming yields tangible benefits for security leaders in any organization. OpenAI’s paper on external red teaming provides a detailed analysis of how the company strives to create specialized external teams that include cybersecurity and subject matter experts. The goal is to see if knowledgeable external teams can defeat models’ security perimeters and find gaps in their security, biases and controls that prompt-based testing couldn’t find. What makes OpenAI’s recent papers noteworthy is how well they define using human-in-the-middle

Read More »