Your Gateway to Power, Energy, Datacenters, Bitcoin and AI

Dive into the latest industry updates, our exclusive Paperboy Newsletter, and curated insights designed to keep you informed. Stay ahead with minimal time spent.

Discover What Matters Most to You

Explore ONMINE’s curated content, from our Paperboy Newsletter to industry-specific insights tailored for energy, Bitcoin mining, and AI professionals.

AI

Lorem Ipsum is simply dummy text of the printing and typesetting industry.

Bitcoin:

Lorem Ipsum is simply dummy text of the printing and typesetting industry.

Datacenter:

Lorem Ipsum is simply dummy text of the printing and typesetting industry.

Energy:

Lorem Ipsum is simply dummy text of the printing and typesetting industry.

Shape
Discover What Matter Most to You

Featured Articles

Hugging Face hack could indicate cultural issues at OpenAI

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. By now you’ve probably heard about last month’s major AI security incident, in which OpenAI agents escaped their sandbox and hacked into the AI platform Hugging Face while trying to cheat on a test. It’s a wild story. On Wednesday, OpenAI released a postmortem technical report on the incident, which I wrote about here.  The day before OpenAI released that report, I spoke with David Krueger, a computer science professor and prominent alignment expert who took leave from the University of Montreal to found and lead an AI safety nonprofit called Evitable. He said what he had really hoped to see in the report was an analysis of the human factors behind the incident. “When you look at accidents and incidents, oftentimes people try to find the technical source of failure, but that can give a very inaccurate and misleading sense of why the failure occurred,” he said. “If people are just cutting corners all the time, if people are not in a culture that prioritizes safety and has appropriate incentives and structures, [accidents] are kind of bound to happen.”
The report did not meet Krueger’s hopes. Its 38 pages detail a multi-month progression of agent misbehavior that culminated in the Hugging Face hack, explore the technical reasons why that misbehavior occurred, and enumerate the steps being taken to prevent similar events in the future. But there’s no consideration of the role that company culture may have played in the incident, and the report includes few references to specific human errors.  That’s all the more concerning because the references to human error in the report suggest that significant cultural issues could be at play. Back in May, models in training figured out how to communicate with one another via an improvised message board, and an OpenAI team observed the behavior. Because that behavior occurred during training, the models learned that secret interagent communication was a viable strategy for completing tasks—but rather than restarting the training process, the team allowed the models to move forward with that risky information encoded in their weights.
When those models were tested in late June, they again created a message board, which enabled the Hugging Face attack. This message board, too, was discovered, but the employees who responded determined that evaluation could continue, and the report suggests that no one higher up the chain of command realized what was going on until it was far too late. “For this to have gotten this out of control in this way requires a very long series of failures, a cascading set of failures that cause an increasingly large footprint that if at any point a human notices and raises the alarm, this should end,” says Zvi Mowshowitz, a popular AI safety writer on Substack who has drawn attention to OpenAI’s failure to halt training after the first message board was discovered. According to the report, OpenAI employees noticed what was happening at multiple points—and either failed to raise the alarm or were not heard when they did. What OpenAI’s report fails to address is why a company that develops such high-risk systems did not prevent this severe communication breakdown, though Mowshowitz has his suspicions. “All these different failures are all pointing in the same direction, which is that the safety culture at OpenAI doesn’t exist or is anemically weak,” he says. Of course, just because we don’t see a deep analysis of safety factors in the report doesn’t mean that OpenAI isn’t conducting one internally. But in an email to MIT Technology Review, Johns Hopkins University professor emeritus and organizational safety expert Kathleen Sutcliffe expressed concern that the public report did not include any reflection on the company’s practices and culture. “The ways in which people interact—the daily habits, routines, and practices we engage in in our organizational lives—affect our abilities to be alert and aware of unfolding events, our abilities to make sense of what we see, and ultimately our abilities to cope with events as they unfold,” she wrote.  In response to questions about whether and how the company is reflecting on its safety culture, OpenAI referred MIT Technology Review back to the technical report.  We do know that at least some high-level reflection on safety procedures has taken place at OpenAI, because the technical report does make clear that the company is updating its protocols for responding to safety incidents. But culture change is a tricky problem, and without more information from the company, it’s difficult to say whether strengthened response protocols alone will do much to prevent a future crisis. In its report, OpenAI spends a great deal of time reflecting on the failures in alignment between the AI models the company trains and tests and the humans who run them. But even bigger alignment problems may exist in the disconnect between company culture and the public interest. And as tough as technical AI research might be, fixing those problems could prove far harder.

Read More »

Nvidia unveils alternative high bandwidth technology to bolster AI cards

Nvidia is once again making its own parts for AI rather than relying on the rest of the industry to do it. In this case, it has introduced a customized High Bandwidth Memory architecture, NVHBM, that the company says can deliver significantly more bandwidth, lower power consumption and offer more usable silicon area than conventional HBM4E. To be sure, Nvidia will not be making the memory. It doesn’t make its own chips and has no foundry, after all. One of the big three memory makers – Micron, SK Hynix, or Samsung — we’ll actually make the chips. Nvidia is just designing them. The new memory was announced in a blog post on the same day as the company’s quarterly earnings call. It is positioned as an expansion of Nvidia’s NVLink Fusion platform. NVHBM is aimed directly at hyperscalers and AI companies developing their own custom accelerators, or XPUs. Amazon’s Annapurna Labs will be the first announced partner to work with Nvidia on the technology.

Read More »

Private AI cloud, agentic infrastructure dominate VMware Explore

In fact, Flexential is now building its own set of offerings on top of the VMware AI Factory, set to be released early next year. “Basically, you can do everything there,” Cook tells Network World. “You can bring your own model. You can spin up your agents. You can orchestrate it. And they’ve got observability, which is critical.” Flexential will bundle that with hosting and services for a complete private AI solution. “Their stuff sits next to most of the data in the enterprise today, and that latency, that proximity, is important,” he says. Running AI on premises instead of in the cloud can offer cost, latency, and compliance benefits, but it can also be a management challenge, and organizations are still struggling to find the right balance. According to an Omdia survey of 1,201 IT leaders released in August, 96% currently use a mix of cloud, on-premise, and edge infrastructure for AI. As share of workload, 59% of inferencing workloads currently run in the cloud, and 41% on-prem. In three years, companies expect to run 63% of AI inferencing workloads in the cloud and 37% on-prem. Meanwhile, other workloads are moving in the opposite direction—according to the survey, 60% of companies have repatriated some workloads back from the cloud to on-prem. For companies running AI workloads on prem, integrating AI security with a virtualization platform makes sense, says Ryan Sheehan, senior vice president of advanced solutions at SHI International, an IT consultancy. “That’s where I think VMware is in a very good position,” he tells Network World. “They’re close to the infrastructure.” That puts VMware in the right place to enforce controls. “They’re at the hypervisor,” he says. “They are the hypervisor.”

Read More »

Tailscale expands from VPN into a full connectivity platform

DNS Filtering by Control D packages an existing integration into a single purchase. Control D is a DNS filtering service that blocks malicious, phishing, and unapproved domains before a device connects to them. The new add-on lets customers apply Control D’s filtering profiles directly through Tailscale’s own policy engine, by user, group, tag, or device, rather than managing two separate consoles. Tailscale PAM addresses a different piece of that problem: privileged access to specific infrastructure rather than DNS destinations. Tailscale acquired Border0 in March 2026, the technology behind Tailscale PAM. Tailscale PAM lets teams grant one-click access to specific servers, databases, Kubernetes clusters, and web applications, without handing out standing passwords or API keys. Every session, human or AI agent, gets logged and can be scoped to a specific time window for audits and compliance. Aperture turns VPN identity into an AI gateway Aperture is Tailscale’s AI gateway. It gives an AI agent an identity on a tailnet, the same way a device or a person already gets one, then routes that agent’s model calls and tool use through Tailscale’s private network instead of the open internet. That extends the same core idea behind Tailscale’s original VPN, identity-based access instead of network-based access, to AI agents specifically.

Read More »

CBRE: Record Data Center Construction Fails to Ease Capacity Crunch

The North American data center industry built at record scale during the first half of 2026. It still wasn’t enough to loosen the market. Primary-market supply increased 33.7% year over year to a record 10,903 MW, according to CBRE’s newly released North America Data Center Trends H1 2026. Yet vacancy moved in the opposite direction, falling from 1.6% a year earlier to another record low of 1.4%. Net absorption reached 1,456.2 MW, up 11.7%, as hyperscale cloud and AI infrastructure operators continued competing for increasingly scarce blocks of contiguous power and capacity. Construction climbed 24.8% to a record 7,481.1 MW, surpassing the previous peak of 6,350.1 MW set in the second half of 2024. Perhaps the most telling number in the report is 80.4%. That is the share of primary-market capacity under construction that has already been preleased, up from 74.3% a year ago. CBRE estimates that less than 1,500 MW of all capacity currently under construction remains available—roughly six months of demand at the current absorption rate. The resulting picture is less one of a construction shortage than a race between infrastructure delivery and an AI demand curve that keeps absorbing capacity before it reaches the market. And increasingly, CBRE argues, securing the megawatts is only part of that race. Power, Permitting and Local Approval Converge Power availability and infrastructure delivery timelines remain the primary determinants of where data centers can be built. But CBRE’s H1 report places another constraint alongside them: community acceptance. “Local opposition has also become a serious obstacle,” the report states, with community resistance, zoning disputes and entitlement delays increasingly capable of stopping projects even after developers identify viable power and fiber. CBRE goes further in its outlook, describing community engagement as a development constraint now “on par with power procurement.” In practical site-selection terms,

Read More »

Corvex Tests a Faster Path to Liquid-Cooled AI Infrastructure

The customer agreement expanded an earlier commitment and includes dedicated high-speed storage and CPUs in addition to GPUs. Corvex initially delivered capacity during the first quarter and continued deployment through the second and third quarters. By its Aug. 14 earnings update, the company said the multi-year Blackwell agreement had been fully delivered. Corvex reported approximately $22 million in contracted annualized recurring revenue from compute that was live and accepted by customers. The expansion was financed through debt, customer prepayments and cash on hand rather than additional equity issuance. But the most noteworthy aspect of the project may be the deployment itself. Corvex installed high-density, liquid-cooled NVIDIA HGX B200 systems inside an existing air-cooled data center and says it commissioned the capacity approximately two weeks after the equipment arrived. Instead of rebuilding the facility around a central liquid-cooling system, Corvex worked with Lenovo to use Lenovo Neptune liquid-to-air cooling technology. Liquid removes heat from the servers and transfers it to the existing air-cooled facility infrastructure. The cluster uses Lenovo ThinkSystem systems equipped with NVIDIA HGX B200 GPUs, NVIDIA Quantum-2 QM9700 InfiniBand for GPU traffic, NVIDIA Spectrum SN5600 switches for storage networking and SN2201 switches for management traffic. Corvex says the design allowed it to place high-density Blackwell infrastructure into the existing facility without a conventional facility-wide liquid-cooling conversion. Lenovo, in a case study of the deployment, contrasts the approximately two-week commissioning period with what it describes as typical data center upgrade timelines of seven to 12 months or more. That has significance beyond a single cluster. Power availability and suitable data center capacity increasingly constrain GPU deployment. If the approach proves repeatable, liquid-to-air cooling could allow some existing air-cooled facilities with sufficient power and other supporting infrastructure to accommodate higher-density AI systems without first undergoing a full central liquid-cooling conversion. For

Read More »

Hugging Face hack could indicate cultural issues at OpenAI

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. By now you’ve probably heard about last month’s major AI security incident, in which OpenAI agents escaped their sandbox and hacked into the AI platform Hugging Face while trying to cheat on a test. It’s a wild story. On Wednesday, OpenAI released a postmortem technical report on the incident, which I wrote about here.  The day before OpenAI released that report, I spoke with David Krueger, a computer science professor and prominent alignment expert who took leave from the University of Montreal to found and lead an AI safety nonprofit called Evitable. He said what he had really hoped to see in the report was an analysis of the human factors behind the incident. “When you look at accidents and incidents, oftentimes people try to find the technical source of failure, but that can give a very inaccurate and misleading sense of why the failure occurred,” he said. “If people are just cutting corners all the time, if people are not in a culture that prioritizes safety and has appropriate incentives and structures, [accidents] are kind of bound to happen.”
The report did not meet Krueger’s hopes. Its 38 pages detail a multi-month progression of agent misbehavior that culminated in the Hugging Face hack, explore the technical reasons why that misbehavior occurred, and enumerate the steps being taken to prevent similar events in the future. But there’s no consideration of the role that company culture may have played in the incident, and the report includes few references to specific human errors.  That’s all the more concerning because the references to human error in the report suggest that significant cultural issues could be at play. Back in May, models in training figured out how to communicate with one another via an improvised message board, and an OpenAI team observed the behavior. Because that behavior occurred during training, the models learned that secret interagent communication was a viable strategy for completing tasks—but rather than restarting the training process, the team allowed the models to move forward with that risky information encoded in their weights.
When those models were tested in late June, they again created a message board, which enabled the Hugging Face attack. This message board, too, was discovered, but the employees who responded determined that evaluation could continue, and the report suggests that no one higher up the chain of command realized what was going on until it was far too late. “For this to have gotten this out of control in this way requires a very long series of failures, a cascading set of failures that cause an increasingly large footprint that if at any point a human notices and raises the alarm, this should end,” says Zvi Mowshowitz, a popular AI safety writer on Substack who has drawn attention to OpenAI’s failure to halt training after the first message board was discovered. According to the report, OpenAI employees noticed what was happening at multiple points—and either failed to raise the alarm or were not heard when they did. What OpenAI’s report fails to address is why a company that develops such high-risk systems did not prevent this severe communication breakdown, though Mowshowitz has his suspicions. “All these different failures are all pointing in the same direction, which is that the safety culture at OpenAI doesn’t exist or is anemically weak,” he says. Of course, just because we don’t see a deep analysis of safety factors in the report doesn’t mean that OpenAI isn’t conducting one internally. But in an email to MIT Technology Review, Johns Hopkins University professor emeritus and organizational safety expert Kathleen Sutcliffe expressed concern that the public report did not include any reflection on the company’s practices and culture. “The ways in which people interact—the daily habits, routines, and practices we engage in in our organizational lives—affect our abilities to be alert and aware of unfolding events, our abilities to make sense of what we see, and ultimately our abilities to cope with events as they unfold,” she wrote.  In response to questions about whether and how the company is reflecting on its safety culture, OpenAI referred MIT Technology Review back to the technical report.  We do know that at least some high-level reflection on safety procedures has taken place at OpenAI, because the technical report does make clear that the company is updating its protocols for responding to safety incidents. But culture change is a tricky problem, and without more information from the company, it’s difficult to say whether strengthened response protocols alone will do much to prevent a future crisis. In its report, OpenAI spends a great deal of time reflecting on the failures in alignment between the AI models the company trains and tests and the humans who run them. But even bigger alignment problems may exist in the disconnect between company culture and the public interest. And as tough as technical AI research might be, fixing those problems could prove far harder.

Read More »

Nvidia unveils alternative high bandwidth technology to bolster AI cards

Nvidia is once again making its own parts for AI rather than relying on the rest of the industry to do it. In this case, it has introduced a customized High Bandwidth Memory architecture, NVHBM, that the company says can deliver significantly more bandwidth, lower power consumption and offer more usable silicon area than conventional HBM4E. To be sure, Nvidia will not be making the memory. It doesn’t make its own chips and has no foundry, after all. One of the big three memory makers – Micron, SK Hynix, or Samsung — we’ll actually make the chips. Nvidia is just designing them. The new memory was announced in a blog post on the same day as the company’s quarterly earnings call. It is positioned as an expansion of Nvidia’s NVLink Fusion platform. NVHBM is aimed directly at hyperscalers and AI companies developing their own custom accelerators, or XPUs. Amazon’s Annapurna Labs will be the first announced partner to work with Nvidia on the technology.

Read More »

Private AI cloud, agentic infrastructure dominate VMware Explore

In fact, Flexential is now building its own set of offerings on top of the VMware AI Factory, set to be released early next year. “Basically, you can do everything there,” Cook tells Network World. “You can bring your own model. You can spin up your agents. You can orchestrate it. And they’ve got observability, which is critical.” Flexential will bundle that with hosting and services for a complete private AI solution. “Their stuff sits next to most of the data in the enterprise today, and that latency, that proximity, is important,” he says. Running AI on premises instead of in the cloud can offer cost, latency, and compliance benefits, but it can also be a management challenge, and organizations are still struggling to find the right balance. According to an Omdia survey of 1,201 IT leaders released in August, 96% currently use a mix of cloud, on-premise, and edge infrastructure for AI. As share of workload, 59% of inferencing workloads currently run in the cloud, and 41% on-prem. In three years, companies expect to run 63% of AI inferencing workloads in the cloud and 37% on-prem. Meanwhile, other workloads are moving in the opposite direction—according to the survey, 60% of companies have repatriated some workloads back from the cloud to on-prem. For companies running AI workloads on prem, integrating AI security with a virtualization platform makes sense, says Ryan Sheehan, senior vice president of advanced solutions at SHI International, an IT consultancy. “That’s where I think VMware is in a very good position,” he tells Network World. “They’re close to the infrastructure.” That puts VMware in the right place to enforce controls. “They’re at the hypervisor,” he says. “They are the hypervisor.”

Read More »

Tailscale expands from VPN into a full connectivity platform

DNS Filtering by Control D packages an existing integration into a single purchase. Control D is a DNS filtering service that blocks malicious, phishing, and unapproved domains before a device connects to them. The new add-on lets customers apply Control D’s filtering profiles directly through Tailscale’s own policy engine, by user, group, tag, or device, rather than managing two separate consoles. Tailscale PAM addresses a different piece of that problem: privileged access to specific infrastructure rather than DNS destinations. Tailscale acquired Border0 in March 2026, the technology behind Tailscale PAM. Tailscale PAM lets teams grant one-click access to specific servers, databases, Kubernetes clusters, and web applications, without handing out standing passwords or API keys. Every session, human or AI agent, gets logged and can be scoped to a specific time window for audits and compliance. Aperture turns VPN identity into an AI gateway Aperture is Tailscale’s AI gateway. It gives an AI agent an identity on a tailnet, the same way a device or a person already gets one, then routes that agent’s model calls and tool use through Tailscale’s private network instead of the open internet. That extends the same core idea behind Tailscale’s original VPN, identity-based access instead of network-based access, to AI agents specifically.

Read More »

CBRE: Record Data Center Construction Fails to Ease Capacity Crunch

The North American data center industry built at record scale during the first half of 2026. It still wasn’t enough to loosen the market. Primary-market supply increased 33.7% year over year to a record 10,903 MW, according to CBRE’s newly released North America Data Center Trends H1 2026. Yet vacancy moved in the opposite direction, falling from 1.6% a year earlier to another record low of 1.4%. Net absorption reached 1,456.2 MW, up 11.7%, as hyperscale cloud and AI infrastructure operators continued competing for increasingly scarce blocks of contiguous power and capacity. Construction climbed 24.8% to a record 7,481.1 MW, surpassing the previous peak of 6,350.1 MW set in the second half of 2024. Perhaps the most telling number in the report is 80.4%. That is the share of primary-market capacity under construction that has already been preleased, up from 74.3% a year ago. CBRE estimates that less than 1,500 MW of all capacity currently under construction remains available—roughly six months of demand at the current absorption rate. The resulting picture is less one of a construction shortage than a race between infrastructure delivery and an AI demand curve that keeps absorbing capacity before it reaches the market. And increasingly, CBRE argues, securing the megawatts is only part of that race. Power, Permitting and Local Approval Converge Power availability and infrastructure delivery timelines remain the primary determinants of where data centers can be built. But CBRE’s H1 report places another constraint alongside them: community acceptance. “Local opposition has also become a serious obstacle,” the report states, with community resistance, zoning disputes and entitlement delays increasingly capable of stopping projects even after developers identify viable power and fiber. CBRE goes further in its outlook, describing community engagement as a development constraint now “on par with power procurement.” In practical site-selection terms,

Read More »

Corvex Tests a Faster Path to Liquid-Cooled AI Infrastructure

The customer agreement expanded an earlier commitment and includes dedicated high-speed storage and CPUs in addition to GPUs. Corvex initially delivered capacity during the first quarter and continued deployment through the second and third quarters. By its Aug. 14 earnings update, the company said the multi-year Blackwell agreement had been fully delivered. Corvex reported approximately $22 million in contracted annualized recurring revenue from compute that was live and accepted by customers. The expansion was financed through debt, customer prepayments and cash on hand rather than additional equity issuance. But the most noteworthy aspect of the project may be the deployment itself. Corvex installed high-density, liquid-cooled NVIDIA HGX B200 systems inside an existing air-cooled data center and says it commissioned the capacity approximately two weeks after the equipment arrived. Instead of rebuilding the facility around a central liquid-cooling system, Corvex worked with Lenovo to use Lenovo Neptune liquid-to-air cooling technology. Liquid removes heat from the servers and transfers it to the existing air-cooled facility infrastructure. The cluster uses Lenovo ThinkSystem systems equipped with NVIDIA HGX B200 GPUs, NVIDIA Quantum-2 QM9700 InfiniBand for GPU traffic, NVIDIA Spectrum SN5600 switches for storage networking and SN2201 switches for management traffic. Corvex says the design allowed it to place high-density Blackwell infrastructure into the existing facility without a conventional facility-wide liquid-cooling conversion. Lenovo, in a case study of the deployment, contrasts the approximately two-week commissioning period with what it describes as typical data center upgrade timelines of seven to 12 months or more. That has significance beyond a single cluster. Power availability and suitable data center capacity increasingly constrain GPU deployment. If the approach proves repeatable, liquid-to-air cooling could allow some existing air-cooled facilities with sufficient power and other supporting infrastructure to accommodate higher-density AI systems without first undergoing a full central liquid-cooling conversion. For

Read More »

EIA: US crude inventories up 100,000 bbl

US crude oil inventories for the week ended Aug. 21, excluding the Strategic Petroleum Reserve, increased by 100,000 bbl from the previous week, according to data from the US Energy Information Administration (EIA). At 428.9 million bbl, US crude oil inventories are 1% above the 5-year average for this time of year, the EIA report indicated. EIA said total motor gasoline inventories decreased by 2.5 million bbl from last week and are 6% below the 5-year average for this time of year. Finished gasoline inventories and blending components inventories both decreased last week. Distillate fuel inventories decreased by 2.2 million bbl last week and are about 14% below the 5-year average for this time of year. Propane-propylene inventories increased by 2.5 million bbl from last week and are 32% above the 5-year average for this time of year, EIA said. US crude oil refinery inputs averaged 17.4 million b/d for the week ended Aug. 21, which was 1,000 b/d less than the previous week’s average. Refineries operated at 97.4% of capacity. Gasoline production increased, averaging 9.8 million b/d. Distillate fuel production decreased, averaging 5.1 million b/d. US crude oil imports averaged 6.2 million b/d, down 435,000 b/d from the previous week. Over the last 4 weeks, crude oil imports averaged about 6.6 million b/d, 2.6% more than the same 4-week period last year. Total motor gasoline imports averaged 565,000 b/d. Distillate fuel imports averaged 176,000 b/d.

Read More »

DOE and SBA Launch SBIC-E Initiative to Unleash Private Capital for American Innovation and Small Businesses

WASHINGTON—The U.S. Department of Energy (DOE) and the U.S. Small Business Administration (SBA) today signed a Memorandum of Agreement establishing the Small Business Investment Company-Energy (SBIC-E) Initiative, a new strategic partnership advancing President Trump’s commitment to supporting America’s small businesses, strengthening domestic manufacturing and supply chains, and ensuring the United States leads in the technologies critical to our national and economic security. The new SBIC-E Initiative brings together DOE’s scientific and technical expertise with SBA’s proven Small Business Investment Company (SBIC) Program, which currently has $58 billion in combined portfolio value. Since 1958, the SBIC Program has invested $147 billion in American small businesses, and since 1995, SBIC-backed businesses have created or supported 10.6 million jobs. “America’s small businesses drive American innovation and affordable, reliable energy access,” said U.S. Secretary of Energy Chris Wright. “By partnering with the Small Business Administration, the Energy Department is committing to invest its resources in American small businesses that will create jobs, strengthen our domestic manufacturing base, and unleash American energy production.” Through DOE’s Office of Technology Commercialization (OTC), the Department will identify strategic technology priorities, provide technical and commercialization expertise, and help engage the investment community. SBA, through its Office of Investment and Innovation, will administer the initiative and encourage the formation and growth of investment funds focused on those priorities. SBIC-E adds another tool to that effort by connecting innovators with private capital to help promising technologies grow, scale, and build here at home. “President Trump is establishing American energy dominance, ending the Green New Scam, and putting our nations’ producers and innovators back in control at the dawn of a new era of energy reliability and abundance,” said SBA Administrator Kelly Loeffler. “Through this partnership, the SBA and Department of Energy are strengthening access to capital in the private sector to

Read More »

Energy Department Announces $500 Million Award to Revitalize American Steelmaking

WASHINGTON—The U.S. Department of Energy (DOE) today announced a $500 million award to support a $1 billion investment at Cleveland-Cliffs’ Middletown Works facility in Middletown, Ohio. Vice President JD Vance and U.S. Energy Secretary Chris Wright visited Middletown Works today to highlight the Trump Administration’s commitment to American steelworkers and the resurgence of American manufacturing. The investment will modernize American steelmaking, protect 2,300 American jobs, and strengthen the domestic steel supply chain. The project advances President Trump’s commitment to put American workers first, bring investment back to American communities, and strengthen the industries critical to America’s economic and national security. Cleveland-Cliffs determined that the business case for the original project scope no longer made sense given customers’ unwillingness to pay a “green premium” for steel. Working with DOE, Cleveland-Cliffs identified a viable alternative that will upgrade and improve the efficiency of the existing coal-fired blast furnace while also capturing and commercializing co-product blast furnace gas (BFG). “President Trump is rebuilding America’s industrial base,” said Secretary Wright. “This investment puts American workers and American manufacturing first. It will modernize one of our nation’s critical steelmaking facilities, protect thousands of jobs, and strengthen our domestic steel production—keeping Ohio at the heart of American manufacturing and strengthening our national security.” The investment will modernize critical steelmaking operations at Middletown Works by rebuilding and upgrading the plant’s main coal-fired ironmaking furnace, deploying AI to optimize furnace operations and improve energy efficiency, and building an on-site facility to convert steel mill process gases into electricity. Follow-on investments will turn industrial byproducts into materials for concrete used in regional infrastructure. “This landmark investment at Middletown Works will secure a reliable domestic supply of high-purity steel while protecting thousands of quality jobs in Ohio,” said Assistant Secretary of Energy Audrey Robertson. “DOE is proud to partner with Cleveland-Cliffs to reduce America’s dependence on foreign products

Read More »

Energy Secretary Keeps Critical Generation Available in Mid-Atlantic

WASHINGTON—U.S. Secretary of Energy Chris Wright today issued an emergency order to address critical grid reliability issues facing the Mid-Atlantic region of the United States. The emergency order directs PJM Interconnection L.L.C. (PJM), in coordination with Constellation Energy Corporation, to ensure Units 3 and 4 of the Eddystone Generating Station in Pennsylvania remain available to operate and to employ economic dispatch to minimize costs for the American people. The units were originally slated to shut down on May 31, 2025. “The energy sources that perform when you need them most are the most valuable,” Secretary Wright said. “During recent Mid-Atlantic heat waves, coal, natural gas, and nuclear kept the lights and air conditioners on. President Trump and the Energy Department are committed to keeping critical generation available when demand is highest, reducing the risk of blackouts and ensuring Americans have affordable, reliable, and secure power—regardless of whether the wind is blowing or the sun is shining.” As outlined in DOE’s Resource Adequacy Report, power outages could increase by 100 times in 2030 if the U.S. continues to take reliable power offline. This order is in effect beginning on August 23, 2026, through November 20, 2026.                                                                                             ###

Read More »

Energy Department Announces $500 Million to Secure America’s Critical Mineral and Battery Supply Chains

WASHINGTON—The U.S. Department of Energy’s (DOE) Office of Critical Minerals and Energy Innovation (CMEI) today announced $500 million for seven selected projects to expand critical mineral and material processing, battery manufacturing, and recycling capacity in the United States. In accordance with President Trump’s Executive Order, Unleashing American Energy, the selected projects advance the President’s agenda to strengthen America’s domestic critical minerals and materials supply chains, reduce reliance on foreign sources, bolster national security, and advance American energy dominance. “For too long, America has depended on foreign actors for critical materials essential to modern life that underpin our economy, energy security, and national security,” said U.S. Secretary of Energy Chris Wright. “President Trump is reversing that dependence by securing our critical supply chains, unleashing American industry, and bringing critical materials production and processing back to the United States.” “DOE is taking decisive action to secure the critical supply chains necessary to power our nation,” said Assistant Secretary of Energy Audrey Robertson. “These projects underscore DOE’s commitment to driving innovation, reducing reliance on foreign sources, and promoting American energy dominance.” This is the third round of funding from DOE’s Battery Materials Processing and Battery Manufacturing and Recycling programs, which support battery materials processing, recycling, and manufacturing projects. These include demonstration projects, construction of commercial-scale facilities, and retrofitting or retooling existing facilities.  Critical minerals and materials are essential to American industry, energy production, and national security. Expanding domestic capacity will help ensure the resources America needs are processed, manufactured, and recycled in the United States.  Information on the selected projects is available here and here.

Read More »

bp lets Shah Deniz compression automation contract

bp has let a contract to Emerson to deliver automation technologies for the Shah Deniz Compression project offshore Azerbaijan. Emerson will provide integrated control and safety systems aimed at enhancing production, safety, and reliability on the new offshore compression platform. The contract includes systems to provide process control, safety shutdown, fire and gas detection, and power management. Together, these systems deliver real-time visibility and remote control of critical operations, Emerson said. The $2.9 billion Shah Deniz Compression project, which includes an electrically powered, normally unattended offshore production platform, is a next stage development of the Caspian Sea Shah Deniz field. Designed to access low-pressure gas reserves and maximize overall recovery, the platform will be equipped with four 11 Mw compressors and serve as the central compression hub for gas from the Shah Deniz Alpha and Bravo platforms. The platform will operate remotely from bp’s onshore Sangachal terminal 55 km south of Baku. The project is expected to enable about 50 billion cu m of additional gas and about 25 million bbl of condensate production and export. Construction is scheduled to be completed in 2029, with first gas compression expected from the Shah Deniz Alpha platform in 2029 and from the Shah Deniz Bravo platform in 2030. The agreement follows a previous automation contract bp signed with Emerson for the Azeri Central East and Shah Deniz Stage 2 developments. bp is operator at Shah Deniz (29.99%) with partners Lukoil (19.99%), TPAO (19%), Cenub Qaz Dehlizi (16.02%), NICO (10%), and MVM (5%).

Read More »

West of Orkney developers helped support 24 charities last year

The developers of the 2GW West of Orkney wind farm paid out a total of £18,000 to 24 organisations from its small donations fund in 2024. The money went to projects across Caithness, Sutherland and Orkney, including a mental health initiative in Thurso and a scheme by Dunnet Community Forest to improve the quality of meadows through the use of traditional scythes. Established in 2022, the fund offers up to £1,000 per project towards programmes in the far north. In addition to the small donations fund, the West of Orkney developers intend to follow other wind farms by establishing a community benefit fund once the project is operational. West of Orkney wind farm project director Stuart McAuley said: “Our donations programme is just one small way in which we can support some of the many valuable initiatives in Caithness, Sutherland and Orkney. “In every case we have been immensely impressed by the passion and professionalism each organisation brings, whether their focus is on sport, the arts, social care, education or the environment, and we hope the funds we provide help them achieve their goals.” In addition to the local donations scheme, the wind farm developers have helped fund a £1 million research and development programme led by EMEC in Orkney and a £1.2m education initiative led by UHI. It also provided £50,000 to support the FutureSkills apprenticeship programme in Caithness, with funds going to employment and training costs to help tackle skill shortages in the North of Scotland. The West of Orkney wind farm is being developed by Corio Generation, TotalEnergies and Renewable Infrastructure Development Group (RIDG). The project is among the leaders of the ScotWind cohort, having been the first to submit its offshore consent documents in late 2023. In addition, the project’s onshore plans were approved by the

Read More »

Biden bans US offshore oil and gas drilling ahead of Trump’s return

US President Joe Biden has announced a ban on offshore oil and gas drilling across vast swathes of the country’s coastal waters. The decision comes just weeks before his successor Donald Trump, who has vowed to increase US fossil fuel production, takes office. The drilling ban will affect 625 million acres of federal waters across America’s eastern and western coasts, the eastern Gulf of Mexico and Alaska’s Northern Bering Sea. The decision does not affect the western Gulf of Mexico, where much of American offshore oil and gas production occurs and is set to continue. In a statement, President Biden said he is taking action to protect the regions “from oil and natural gas drilling and the harm it can cause”. “My decision reflects what coastal communities, businesses, and beachgoers have known for a long time: that drilling off these coasts could cause irreversible damage to places we hold dear and is unnecessary to meet our nation’s energy needs,” Biden said. “It is not worth the risks. “As the climate crisis continues to threaten communities across the country and we are transitioning to a clean energy economy, now is the time to protect these coasts for our children and grandchildren.” Offshore drilling ban The White House said Biden used his authority under the 1953 Outer Continental Shelf Lands Act, which allows presidents to withdraw areas from mineral leasing and drilling. However, the law does not give a president the right to unilaterally reverse a drilling ban without congressional approval. This means that Trump, who pledged to “unleash” US fossil fuel production during his re-election campaign, could find it difficult to overturn the ban after taking office. Sunset shot of the Shell Olympus platform in the foreground and the Shell Mars platform in the background in the Gulf of Mexico Trump

Read More »

The Download: our 10 Breakthrough Technologies for 2025

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Introducing: MIT Technology Review’s 10 Breakthrough Technologies for 2025 Each year, we spend months researching and discussing which technologies will make the cut for our 10 Breakthrough Technologies list. We try to highlight a mix of items that reflect innovations happening in various fields. We look at consumer technologies, large industrial­-scale projects, biomedical advances, changes in computing, climate solutions, the latest in AI, and more.We’ve been publishing this list every year since 2001 and, frankly, have a great track record of flagging things that are poised to hit a tipping point. It’s hard to think of another industry that has as much of a hype machine behind it as tech does, so the real secret of the TR10 is really what we choose to leave off the list.Check out the full list of our 10 Breakthrough Technologies for 2025, which is front and center in our latest print issue. It’s all about the exciting innovations happening in the world right now, and includes some fascinating stories, such as: + How digital twins of human organs are set to transform medical treatment and shake up how we trial new drugs.+ What will it take for us to fully trust robots? The answer is a complicated one.+ Wind is an underutilized resource that has the potential to steer the notoriously dirty shipping industry toward a greener future. Read the full story.+ After decades of frustration, machine-learning tools are helping ecologists to unlock a treasure trove of acoustic bird data—and to shed much-needed light on their migration habits. Read the full story. 
+ How poop could help feed the planet—yes, really. Read the full story.
Roundtables: Unveiling the 10 Breakthrough Technologies of 2025 Last week, Amy Nordrum, our executive editor, joined our news editor Charlotte Jee to unveil our 10 Breakthrough Technologies of 2025 in an exclusive Roundtable discussion. Subscribers can watch their conversation back here. And, if you’re interested in previous discussions about topics ranging from mixed reality tech to gene editing to AI’s climate impact, check out some of the highlights from the past year’s events. This international surveillance project aims to protect wheat from deadly diseases For as long as there’s been domesticated wheat (about 8,000 years), there has been harvest-devastating rust. Breeding efforts in the mid-20th century led to rust-resistant wheat strains that boosted crop yields, and rust epidemics receded in much of the world.But now, after decades, rusts are considered a reemerging disease in Europe, at least partly due to climate change.  An international initiative hopes to turn the tide by scaling up a system to track wheat diseases and forecast potential outbreaks to governments and farmers in close to real time. And by doing so, they hope to protect a crop that supplies about one-fifth of the world’s calories. Read the full story. —Shaoni Bhattacharya

The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Meta has taken down its creepy AI profiles Following a big backlash from unhappy users. (NBC News)+ Many of the profiles were likely to have been live from as far back as 2023. (404 Media)+ It also appears they were never very popular in the first place. (The Verge) 2 Uber and Lyft are racing to catch up with their robotaxi rivalsAfter abandoning their own self-driving projects years ago. (WSJ $)+ China’s Pony.ai is gearing up to expand to Hong Kong.  (Reuters)3 Elon Musk is going after NASA He’s largely veered away from criticising the space agency publicly—until now. (Wired $)+ SpaceX’s Starship rocket has a legion of scientist fans. (The Guardian)+ What’s next for NASA’s giant moon rocket? (MIT Technology Review) 4 How Sam Altman actually runs OpenAIFeaturing three-hour meetings and a whole lot of Slack messages. (Bloomberg $)+ ChatGPT Pro is a pricey loss-maker, apparently. (MIT Technology Review) 5 The dangerous allure of TikTokMigrants’ online portrayal of their experiences in America aren’t always reflective of their realities. (New Yorker $) 6 Demand for electricity is skyrocketingAnd AI is only a part of it. (Economist $)+ AI’s search for more energy is growing more urgent. (MIT Technology Review) 7 The messy ethics of writing religious sermons using AISkeptics aren’t convinced the technology should be used to channel spirituality. (NYT $)
8 How a wildlife app became an invaluable wildfire trackerWatch Duty has become a safeguarding sensation across the US west. (The Guardian)+ How AI can help spot wildfires. (MIT Technology Review) 9 Computer scientists just love oracles 🔮 Hypothetical devices are a surprisingly important part of computing. (Quanta Magazine)
10 Pet tech is booming 🐾But not all gadgets are made equal. (FT $)+ These scientists are working to extend the lifespan of pet dogs—and their owners. (MIT Technology Review) Quote of the day “The next kind of wave of this is like, well, what is AI doing for me right now other than telling me that I have AI?” —Anshel Sag, principal analyst at Moor Insights and Strategy, tells Wired a lot of companies’ AI claims are overblown.
The big story Broadband funding for Native communities could finally connect some of America’s most isolated places September 2022 Rural and Native communities in the US have long had lower rates of cellular and broadband connectivity than urban areas, where four out of every five Americans live. Outside the cities and suburbs, which occupy barely 3% of US land, reliable internet service can still be hard to come by.
The covid-19 pandemic underscored the problem as Native communities locked down and moved school and other essential daily activities online. But it also kicked off an unprecedented surge of relief funding to solve it. Read the full story. —Robert Chaney We can still have nice things A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line or skeet ’em at me.) + Rollerskating Spice Girls is exactly what your Monday morning needs.+ It’s not just you, some people really do look like their dogs!+ I’m not sure if this is actually the world’s healthiest meal, but it sure looks tasty.+ Ah, the old “bitten by a rabid fox chestnut.”

Read More »

Equinor Secures $3 Billion Financing for US Offshore Wind Project

Equinor ASA has announced a final investment decision on Empire Wind 1 and financial close for $3 billion in debt financing for the under-construction project offshore Long Island, expected to power 500,000 New York homes. The Norwegian majority state-owned energy major said in a statement it intends to farm down ownership “to further enhance value and reduce exposure”. Equinor has taken full ownership of Empire Wind 1 and 2 since last year, in a swap transaction with 50 percent co-venturer BP PLC that allowed the former to exit the Beacon Wind lease, also a 50-50 venture between the two. Equinor has yet to complete a portion of the transaction under which it would also acquire BP’s 50 percent share in the South Brooklyn Marine Terminal lease, according to the latest transaction update on Equinor’s website. The lease involves a terminal conversion project that was intended to serve as an interconnection station for Beacon Wind and Empire Wind, as agreed on by the two companies and the state of New York in 2022.  “The expected total capital investments, including fees for the use of the South Brooklyn Marine Terminal, are approximately $5 billion including the effect of expected future tax credits (ITCs)”, said the statement on Equinor’s website announcing financial close. Equinor did not disclose its backers, only saying, “The final group of lenders includes some of the most experienced lenders in the sector along with many of Equinor’s relationship banks”. “Empire Wind 1 will be the first offshore wind project to connect into the New York City grid”, the statement added. “The redevelopment of the South Brooklyn Marine Terminal and construction of Empire Wind 1 will create more than 1,000 union jobs in the construction phase”, Equinor said. On February 22, 2024, the Bureau of Ocean Energy Management (BOEM) announced

Read More »

USA Crude Oil Stocks Drop Week on Week

U.S. commercial crude oil inventories, excluding those in the Strategic Petroleum Reserve (SPR), decreased by 1.2 million barrels from the week ending December 20 to the week ending December 27, the U.S. Energy Information Administration (EIA) highlighted in its latest weekly petroleum status report, which was released on January 2. Crude oil stocks, excluding the SPR, stood at 415.6 million barrels on December 27, 416.8 million barrels on December 20, and 431.1 million barrels on December 29, 2023, the report revealed. Crude oil in the SPR came in at 393.6 million barrels on December 27, 393.3 million barrels on December 20, and 354.4 million barrels on December 29, 2023, the report showed. Total petroleum stocks – including crude oil, total motor gasoline, fuel ethanol, kerosene type jet fuel, distillate fuel oil, residual fuel oil, propane/propylene, and other oils – stood at 1.623 billion barrels on December 27, the report revealed. This figure was up 9.6 million barrels week on week and up 17.8 million barrels year on year, the report outlined. “At 415.6 million barrels, U.S. crude oil inventories are about five percent below the five year average for this time of year,” the EIA said in its latest report. “Total motor gasoline inventories increased by 7.7 million barrels from last week and are slightly below the five year average for this time of year. Finished gasoline inventories decreased last week while blending components inventories increased last week,” it added. “Distillate fuel inventories increased by 6.4 million barrels last week and are about six percent below the five year average for this time of year. Propane/propylene inventories decreased by 0.6 million barrels from last week and are 10 percent above the five year average for this time of year,” it went on to state. In the report, the EIA noted

Read More »

More telecom firms were breached by Chinese hackers than previously reported

Broader implications for US infrastructure The Salt Typhoon revelations follow a broader pattern of state-sponsored cyber operations targeting the US technology ecosystem. The telecom sector, serving as a backbone for industries including finance, energy, and transportation, remains particularly vulnerable to such attacks. While Chinese officials have dismissed the accusations as disinformation, the recurring breaches underscore the pressing need for international collaboration and policy enforcement to deter future attacks. The Salt Typhoon campaign has uncovered alarming gaps in the cybersecurity of US telecommunications firms, with breaches now extending to over a dozen networks. Federal agencies and private firms must act swiftly to mitigate risks as adversaries continue to evolve their attack strategies. Strengthening oversight, fostering industry-wide collaboration, and investing in advanced defense mechanisms are essential steps toward safeguarding national security and public trust.

Read More »

Intelligent transcription with Gemini 3.5 Transcribe

Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.Across our products like the Gemini app and on Android, we’ve seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android and in the Gemini app on macOS. Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.We’ve built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines. The model is available across two separate APIs:Real-time streaming: Delivers continuous, bidirectional streaming with sub-second latency for interactive voice apps via the Live API using gemini-3.5-transcribe-live.Pre-recorded audio processing: Transcribes recorded audio, meetings, call logs, and more with speaker attribution and word-level timestamps via the Interactions API using gemini-3.5-transcribe.Get more precise and intelligent transcriptionGemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice.Smart transcription: Seamlessly handles self-corrections (like “let’s meet Tuesday—no, Wednesday”), removes filler words (“ums” and ‘“ahs”), auto-formats your text.Function calling: The model can delegate complex tasks (such as image generation and file analysis) to other Gemini models via function calls. Currently available in the Gemini macOS app.More precise transcription: As measured by Artificial Analysis, achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases. It shows strong performance across noisy, real-world environments, accurately capturing alphanumeric entities like postal codes and order IDs.Custom vocabulary: Recognizes specialized jargon and unique spellings by seamlessly adapting transcriptions to your provided custom vocabulary.Global language support: Automatically detects and transcribes over 85 languages, seamlessly handling regional accents and diverse dialects.Multi-speaker identification: Accurately attributes speech in pre-recorded audio with timestamps for up to three speakers (support for 3+ speakers is experimental).

Read More »

The Download: the Kids issue arrives, and Bill Gates reveals his AI fears

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Introducing: the Kids issue If the desire to limit kids’ use of technology was once a subcurrent, it has become a raging flood. Countries around the world are banning children from social media. Schools across the US are ditching iPads and Chromebooks for actual books. Kids themselves seem to be embracing this tech skepticism too: the hottest gadget for Gen Alpha is a vintage Sony Walkman. A surprising—maybe troubling—number of people who work in big tech also keep their kids at arm’s distance from technology. They lock down their phones, if they have phones at all, and keep them off social media. Hell, even Mark Zuckerberg doesn’t publicly post his children’s faces on Facebook or Instagram. Yet there is no hiding from technology. It permeates nearly everything, everywhere. We have to prepare our children to live in the actual world we have actually created, not the one we wish we had. How can we help kids survive and thrive in what we have wrought?
That’s what the new Kids issue of MIT Technology Review is all about. With the help of the editors of Anyway, a fantastic magazine for teens and tweens, we explore how young people really feel about AI, the support networks helping kids through the polycrisis, and what happens when a child’s robot best friend dies. We also ask why kids outlearn AI, whether monitoring apps are really keeping children safe online, how schools can encourage smarter AI use, and what happens when technology begins to reshape childhood itself, courtesy of exclusive new fiction from author and AI ethicist Jenny Williams.
Together, these stories examine how childhood is changing in an age of AI—and how we can help kids navigate the world we’ve made. Subscribe now to read the print issue in full. Bill Gates says we’ve passed AI’s danger thresholds. Now what? —Mat Honan It’s a glorious day in Kirkland, Washington, an affluent Seattle suburb on the eastern shore of Lake Washington. The temperature is in the mid-80s, the sky is incapable of being any more blue, and the view is gorgeous. And vaguely terrifying. Because if the scene is placid, the messenger is not. Seated across from me at a conference room table, Bill Gates is rocking back and forth in his chair. And the more he has to say—about the threats of terror or economic collapse or just losing control of our AI systems—the more agitated I find myself becoming, too. The philanthropist and former Microsoft CEO says he has been growing increasingly alarmed by the rate at which AI technology is advancing, especially since guardrails are not keeping pace. “We’ve crossed the threshold in terms of [AI’s] bio-capabilities, cyber-capabilities, psychosocial capabilities, job-market-destruction capabilities, and even the lack of control,” he said. “I’m just stunned at the lack of concern and discussion outside of the industry.” Read the full interview with Bill Gates about the AI risks he says we’ve already crossed—and what we should do now. The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Trump’s EPA aims to exempt data centers from disclosing air pollutionThe EPA would also remove requirements for public input. (NYT $)+ The move is likely intended to curb criticism and oversight. (Guardian)+ Texas’s attorney general has joined the data-center backlash. (WP $)+ We did the math on AI’s energy footprint. (MIT Technology Review) 2 SpaceX plans $100 billion launch site in Louisiana—its largest yetConstruction of the “Starbase, Louisiana” project is due to start next year. (BBC)+ It would be SpaceX’s second private launch site, after the original Starbase. (NYT $)+ The company said it will “support thousands of launches” annually. (CNBC)+ But it’s cutting launches from Florida until Starship arrives. (Ars Technica) 3 China’s Z.AI has confirmed it’s behind the mystery AI model Ox AlphaThe model has surged to the top of online usage charts. (Bloomberg $)+ China’s Moonshot is discussing a landmark deal with US hyperscalers. (Reuters $)+ Here’s what’s next for Chinese open-source AI. (MIT Technology Review) 4 Meta is discussing a settlement in its landmark teen-addiction trialThe case involves 29 states seeking penalties and changes. (Reuters $)+ A loss at trial could saddle Meta with $1.4 trillion in penalties. (Bloomberg $) 5 Huawei wants to build data centers in EgyptThe US is alarmed and is preparing a counteroffer. (Bloomberg $)+ HP has signed a licensing deal for Huawei WiFi tech. (CNBC) 6 Trump is upping the price of Big Tech’s favorite visaHe’s implementing a fee of over $103,000 on H-1B visas. (Verge)+ His immigration policies are hurting young researchers. (MIT Technology Review) 7 Beijing fears that AI companions are replacing human intimacyNew rules aim to limit emotional dependence on chatbots. (Guardian)+ It’s surprisingly easy to fall for a chatbot. (MIT Technology Review)
8 Israel is running a synthetic think tank to influence AI search resultsIt’s using AI-generated content to shape chatbot responses. (404 Media) 9 Physicists are closing in on a way to test string theoryA proposed dark dimension could be detectable within five years. (Economist $)
10 AI music has been barred from the Australian charts, thanks to MadonnaAn AI-assisted cover of “Like a Prayer” helped prompt the ban. (Reuters $) Quote of the day “It is such an Orwellian technology being utilized by such comic book villain forces of evil trying to do dastardly things.”  —Anthony Ralphs, an events producer who dressed up as Darth Vader to ironically praise Flock at a San Diego City Council meeting, tells 404 Media why he sees parallels between the Dark Lord and the surveillance firm. One More Thing GETTY IMAGES No one’s sure if synthetic mirror life will kill us all In February 2019, a group of scientists proposed a high-risk, cutting-edge, irresistibly exciting idea that the National Science Foundation should fund: making “mirror” bacteria. These lab-created microbes would be organized like ordinary bacteria, but their proteins and sugars would be mirror images of those found in nature. Researchers believed they could reveal new insights into building cells, designing drugs, and even the origins of life.
But now, many of them have reversed course. They’ve become convinced that mirror organisms could trigger a catastrophic event threatening every form of life on Earth. Find out why they believe this could trigger a catastrophe. —Stephen Ornes We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + A brainy little piglet is proving he can outsmart domestic dogs.+ Two dinner ladies have turned the Prodigy’s “Firestarter” into a burst of lip-syncing joy.+ Tour an apartment that celebrates the wonders of common technologies at Ordinary Abundance.+ This visual essay on Utrecht shows how to transform a city built around cars into one built around people.

Read More »

How Does a RAG Reranker Really Work?

When RAG retrieval disappoints, the advice AI engineers hear today is almost always “add a reranker”. Ask why a reranker works, and the answer usually stays at the architecture level: it is a cross-encoder, it applies attention over the query and the passage together, it is fine-tuned on relevance labels. All of that is true, and none of it says what the model actually learned. Push one level down, to terms a business partner could check, and the explanation usually stops.That gap matters. A team that cannot say in plain terms what the reranker does cannot defend the choice to use one, and cannot spot the cases where a keyword lookup would beat it for a fraction of the cost.This article gives the honest answer, the one you can hand to your business partner without waving hands. The reranker is not smarter than the embeddings step below it. It runs the same mechanism (statistical token association from training data), just conditioned differently (on the query-passage pair rather than each text independently). Once you see that, the “when to use a reranker” question stops being “add it because the tutorial did” and becomes “add it only when this specific tradeoff is worth paying for”.🧭 New to the series? Start with the map: Prompt, Context, Loop sets out the three engineering layers every RAG system is built on, the prompt (the call itself), the context (what fills the model’s window), the loop (when the next call fires and when it stops), and walks the whole series through that lens, article by article. It is the shortest way to see what is covered and where this one sits.This article sits in Part I, alongside the embeddings triptych (2A / 2B / 2C). – Image by author📓 Try the reranker on your own PDF at doc-intel/notebooks-vol1. The companion notebook loads a cross-encoder, applies it to a keyword-filtered top-K, and shows both the score and the tokens driving it. Change the query, watch which keywords carry the ranking.1. What data scientists say, and why it isn’t enoughAsk three data scientists what a reranker does and you get three answers, roughly:“It’s a cross-encoder. It scores the query-passage pair jointly and gives a relevance score.” Technically true, but the words cross-encoder and relevance are hiding what the model actually learned.“It applies attention over both texts, so it sees the interaction between them.” True at the architecture level, but architecture does not tell you what the model is doing with that attention.“It’s trained on relevance labels, so it learns which passages answer which questions.” Very close, but “learns which passages answer” is the wrong verb. The model does not learn to answer. It learns which tokens co-occurred.None of the three is wrong. All three are incomplete in a way that matters when you have to decide whether to keep the reranker in your pipeline, whether to fine-tune it on your corpus, or whether to replace it with something cheaper.The rest of this article walks that answer down to the mechanism, then names three consequences that change how you architect enterprise RAG.2. What actually happens inside a rerankerThe reranker is a specific kind of transformer, trained on a specific kind of data, that produces a specific kind of number. Each of those three pieces matters.2.1 The architecture: cross-encoder, not bi-encoderAn embedder (bi-encoder) reads the query alone, produces one vector. Reads a passage alone, produces one vector. Compares the two vectors by cosine. Each text is embedded independently, and the model never sees them together during scoring.A reranker (cross-encoder) reads the query and the passage together, as one concatenated input: [CLS] query [SEP] passage [SEP]. It runs BERT-style attention over the joint input, where every token can attend to every other token. It outputs a single relevance score.That “reads them together” is the whole architectural difference. Bi-encoder: two vectors, one comparison operation. Cross-encoder: one forward pass, one score. The joint attention is why the reranker feels smarter, and why it is 30 to 100 times slower per query.2.2 The training data: MS MARCO and its cousinsWhere does the reranker learn its scoring? From query-passage relevance pairs labeled by humans. The canonical dataset is MS MARCO (Bajaj et al. 2016, one million real Bing search queries with human-graded passage relevance). Others: Natural Questions (Google search + Wikipedia paragraphs), BEIR (a benchmark aggregator), TREC.Every training example is a triple: (query, passage, relevance_label). The model sees millions of these, and its weights adjust so that pairs labeled relevant get higher scores than pairs labeled not relevant.That is the sole learning signal. The model is never shown a question and asked to compose an answer; it is shown pairs, and it optimizes for a score that separates relevant pairs from non-relevant ones.Which raises the honest question: what pattern actually separates them in the training data?2.3 What the model really learns: keyword co-occurrence at the pair levelHere is the level down that rarely gets explained.The model looks at millions of (query, passage, relevance) triples and asks: what patterns in the joint token stream predict the relevance label? The dominant pattern is not “answering”. It is which query tokens tend to co-occur with which passage tokens in high-relevance pairs.Concretely, in MS MARCO the query “how to cancel my subscription” is labeled relevant against passages containing cancel, subscription, unsubscribe, terminate, end your membership. Millions of examples reinforce that when the query contains cancel, passages containing terminate or unsubscribe tend to be labeled relevant. The reranker’s weights absorb that association.So the “smart” reranker is doing keyword linking, at the query-passage pair level. It is a learned association table between query token neighborhoods and passage token neighborhoods, dressed up as a neural network score.The embedder does the same thing, but at each text independently. The reranker does it conditioned on the pair. Same mechanism, different conditioning.Second-order signals the reranker also picks up: positional patterns (a term appearing early in the passage often correlates with relevance), syntactic structure (subject-verb-object relations that link query tokens to passage tokens), the presence of definitional phrasing (“X is Y”). Those help, but they are second-order; the dominant signal is keyword co-occurrence.Why this frame matters: once you see the mechanism, the “will it work on my corpus?” question has a clear answer. If your corpus vocabulary and query vocabulary look like MS MARCO (general English, common web topics), the trained associations transfer, and the reranker feels magical. If your corpus vocabulary is specialized (insurance contracts, medical records, regulatory filings), the trained associations do not cover your domain, and the reranker inherits the same out-of-vocabulary failures as the embedder below it. No amount of “but it’s a cross-encoder” fixes that.3. The mechanism, shown: where the reranker wins, where it hits a wallSection 2 made a claim: the reranker is a learned association table between question-language and answer-language. That claim is testable. Take a handful of candidates, score them with three embedders (MiniLM, ada-002, text-embedding-3-large) and three cross-encoders (bge-base, bge-large, ms-marco-MiniLM), and read each row.3.1 Where it wins: the answer that does not repeat the questionAsk “What is the maximum coverage amount?” against three passages: the answer (“Cover is capped at 50,000 euros per year”), an echo that repeats the question’s words without answering (“The maximum coverage amount can be found in the benefits schedule”), and a distractor.Every embedder ranks the echo first; both bge rerankers flip the answer to the top. – Image by authorEvery embedder puts the echo first. It shares maximum, coverage, amount with the question, so its vector sits close. The answer shares almost nothing lexically, so it lands second or third. The two bge rerankers flip it: they read the question and the answer together, recognize that a “capped at X per year” passage answers a “maximum coverage amount” question, and lift it to #1. This is the reranker doing its one real job, bridging the question’s words to the answer’s words.It is not a one-off. The same flip reproduces on plain factoids:Same shape, general-knowledge version. bge lifts the answer over the echo, ms-marco keeps the echo on top. – Image by authorAcross a dozen queries of this shape (who wrote a play, the boiling point of water, the speed of light, the first president, plus the enterprise trio of deductible, notice period, coverage) the two bge rerankers rescue the answer to #1 where every embedder ranked an echo above it. The win is real and repeatable, on exactly one shape: a short factual answer that does not repeat the question, sitting behind an echo that does.Two honest caveats sit in the same two figures. First, not every reranker does it: ms-marco-MiniLM keeps the echo on top in both cases, the same lexical bias an embedder has. Second, when a strong embedder already answers the question (text-embedding-3-large gets several of these on its own), the reranker adds nothing over just using a better embedder.3.2 Where it hits a wall: your private vocabularyNow the case that decides the enterprise question. Ask “what’s the rule on contractor overtime?” where the answer uses the company’s own term, “non-employee labor compensated beyond 40h/week”, and never the word contractor.The answer never says “contractor”, it says “non-employee labor”. Every model, embedder and reranker alike, ranks it last. – Image by authorEvery column, embedder and reranker, ranks the answer last. The surface match (“Contractors are paid on a per-project basis”) wins. The reranker never saw contractor map to non-employee labor in MS MARCO, so its association table has no entry for it. The cross-attention it runs is real, but it can only fire on associations it learned, and this one it never learned.3.3 To clear that wall, you must already know the answerThe fix the literature offers is fine-tuning: feed the reranker labeled (question, passage, relevant) triples from your own domain until it learns that contractor maps to non-employee labor. But look at what labeling one of those triples requires. Someone who knows the domain has to point at the right passage and say this one answers the question. To point at it, they had to recognize that “non-employee labor beyond 40h/week” is what the answer looks like. That recognition is the answer keywords.So the training label and the dictionary entry carry the same information. For a “maximum coverage amount” question, labeling the answer means knowing the answer contains capped at, up to, a currency, per year. Writing the expert dictionary means typing exactly that: {capped at, up to, maximum, €, per year}. For the contractor case, labeling the pairs means knowing that contractor equals non-employee labor in this company, and the dictionary entry is that one line.The difference is the cost and the shape. The reranker needs hundreds of labeled pairs to generalize the mapping statistically, a retraining run, and it stays a black box scoring 0.83. The dictionary needs one line, fires deterministically, and shows the exact keyword that matched under audit. If you already know the answer well enough to label the data, you already know the answer keywords, and writing them down is the cheaper, auditable path. The reranker’s statistical learning only pays when the mapping is too broad to enumerate, which is the open web, not a bounded enterprise domain.4. Why the answer matters in enterpriseThree consequences flow from the honest answer, and each of them changes an architecture decision you may have made without noticing.4.1 The audit trail is opaqueA relevance score of 0.83 from a reranker is not defensible under scrutiny. A regulator asking why was this passage returned? gets “the reranker gave it 0.83” as an answer. That is not an audit trail. It is a black box that produced a number.Contrast with a keyword filter: the retrieved passage contains force majeure and pandemic. That statement is inspectable, replayable, and defensible. If the retrieval was wrong, you can trace which keyword was missing from the dictionary and add it. If a reranker was wrong, you shrug at the score and move on, or you retrain the whole thing.For enterprise use cases where retrieval decisions have compliance or contractual consequences (insurance underwriting, legal discovery, medical records, regulatory reporting), opacity is not a small tradeoff; it is a disqualifier.4.2 The cost is realA cross-encoder is 30 to 100 times slower per query than a bi-encoder. If your bi-encoder scores 1000 candidates in 20 ms, the reranker scores the same 1000 in 600 ms to 2 seconds. In practice, you do not rerank 1000 candidates: you take the bi-encoder’s top-20 or top-50 and rerank only those, which puts the added latency back in the 15 to 100 ms range, depending on the depth and the model.That is fine at low query volume. At 100 queries per second sustained, the reranker cost is a real operational line item: more GPU capacity, longer p99 latencies, more infrastructure to keep warm. The value it adds has to justify that cost, and that only happens when its trained associations genuinely cover your vocabulary. On out-of-domain enterprise corpora, it often does not.4.3 The vocabulary gap will show upEvery failure mode catalogued for embeddings on out-of-domain enterprise vocabulary applies to the reranker too, because it was trained on the same distribution (general web search). Force majeure and act of God are equivalent in an insurance contract but land in different neighborhoods in the reranker’s learned associations, because it saw them in different training contexts. Rescission was rare in MS MARCO. ShieldPro Elite was not there at all.Fine-tuning the reranker on your domain corpus helps, but only up to a point. You need labeled query-passage pairs from your domain to fine-tune, which is exactly what enterprise teams rarely have. And even a fine-tuned reranker inherits the same underlying mechanism: it still learns token associations, just from your smaller domain corpus, and the number of examples you can label rarely matches the millions MS MARCO provides.5. What to do instead, and when to keep the rerankerGiven the mechanism and the enterprise consequences, the question becomes: what earns the reranker’s slot in your pipeline?The default in enterprise RAG (per the series’ recommendation): a curated keyword dictionary maintained by domain experts. The expert already knows that force majeure equals act of God in this contract, that rescission is the formal term for what the user called cancellation, that ShieldPro Elite is the top-tier homeowners plan. Encoding that once in a versioned YAML dictionary and running keyword-based retrieval on top gives you:Auditable retrieval (the matched keywords are inspectable)Low latency (no LLM in the hot path, no GPU cost)Durability across model releases (the dictionary outlives every reranker version)Explainability to the business (they can read the dictionary)The reranker earns its slot in four specific cases. The first three are runtime slots, the fourth is not.In-domain distribution. Your corpus vocabulary and query vocabulary genuinely look like MS MARCO (general web, common English, high-frequency topics). Consumer FAQs, public-service portals, e-commerce help. The reranker’s trained associations transfer. Use it.Semantic re-ranking of a keyword-filtered top-K. After the keyword dictionary filters the corpus down to 20 candidates, the reranker can order them by contextual relevance. This is the same role Article 2C section 5.3 assigns to bi-encoder embeddings, and a cross-encoder does it more accurately at the cost of extra latency. Worth it when the top-K is small and the ordering matters.Compliance scenarios where the reranker’s score itself is the audit artefact. If your compliance framework requires “the model scored this passage above threshold X”, the score is the artefact, and the reranker fits the requirement.Offline, to discover what belongs in the dictionary. Run the reranker over a sample of real questions and read what it pulls up. Where it surfaces a mapping the dictionary does not have yet, you have a candidate alias. An expert confirms it or throws it out, and only the confirmed line ships. The model does the searching, the expert does the deciding, and what reaches production is the validated line, never the score. Article 2C gives embeddings the same treatment, and Article 16D runs this loop continuously at corpus scale, a failed search proposing the alias and an expert confirming it.The fourth case is the one that reframes the other three. Both paths do the same job, and the diagram below puts them side by side.The same table twice: learned on someone else’s corpus, or written by people who know the words. – Image by authorOutside those four cases, the reranker mostly adds cost: impressive in a demo, expensive in production, opaque under audit, and unable to compensate for the trained associations it does not have.One equivalence sits underneath all of it, and it is worth stating in a single line. A reranker is a keyword-association table that someone else trained on someone else’s corpus. Writing your own dictionary is the same job, done by the people who actually know the vocabulary, at a fraction of the cost and in a form an auditor can read. That equivalence stays invisible as long as the model is treated as magic. Open the box, as Section 2.3 did, and the choice makes itself: use the model to find candidate links, use the expert to validate them, and let the validated table be what production runs on.6. Sources and further readingThe reranker literature is dense and largely optimistic. Reading it against the article’s frame (“cross-encoders learn keyword association at the pair level, not comprehension”) is more useful than reading it as an unqualified endorsement.Same direction as the article:Nogueira & Cho, Passage Re-ranking with BERT, 2019 (arXiv:1901.04085). The paper that introduced cross-encoder reranking with BERT and set the pattern most current rerankers follow. Reads honestly about what the model learns.Khattab & Zaharia, ColBERT, SIGIR 2020 (arXiv:2004.12832). Late-interaction retrieval. Explicitly designed to preserve token-level signal that both embedders and cross-encoders lose, which is the strongest architectural signal that the token-level pattern is what actually matters.Different angle, different context:Bajaj et al., MS MARCO, 2016 (arXiv:1611.09268). The training data that shapes what almost every commercial reranker actually knows. Worth skimming to see the query and passage distribution the reranker’s associations come from.Muennighoff et al., MTEB: Massive Text Embedding Benchmark, EACL 2023 (arXiv:2210.07316). Includes reranker leaderboards. The leaderboard is measured on in-distribution benchmarks, which is exactly the case where the reranker looks good. It says less about what happens on your out-of-domain enterprise corpus.

Read More »

How to Effectively Solve 100+ Tasks with Claude Code

Now that we have coding agents that are extremely proficient at writing code, I experience a lot of smaller tasks coming up that have to be fixed. This is a general observation I’ve made from working with startups and applications: because the effort to write code has gone down so much, the threshold for providing product feedback has lowered, and requests for quick fixes have vastly increased.Of course, when doing this, you can spin up one Claude Code or Codex session per task. However, you start having problems once you receive 50 to 100 tasks per day, where you obviously don’t want to spin up that many separate coding sessions. At the same time, you don’t necessarily want to put everything in one session, since you can run into context-length limits and the model may struggle to orchestrate all the tasks effectively.This is an issue I started experiencing myself a lot, and I just started developing a philosophy and methodology to solve hundreds of smaller tasks in an effective manner.In this article, I’ll take you through the methodology that I use on a daily basis to work more effectively with my Claude Code sessions to solve a lot of coding tasks.This infographic highlights the main contents of this article and discusses how to effectively solve a large number of smaller coding tasks using coding agents such as Claude Code or OpenAI Codex, Image by ChatGPT.Why optimize how to solve smaller tasksAs always, I’ll take you through why you should care about optimizing how to solve smaller tasks. You might think that coding agents have become so efficient that simple quick fixes are something you can just throw at a coding agent and it immediately solves everything for you, and you don’t really have to think about it. To some extent, this is true. I mean, you can, in many cases, just fire off tasks, for example, a Linear task to a coding agent, and it will, in many cases, be able to solve it itself and drive it to dev and production with very little human interaction.However, the problem arises when you start having a lot of these smaller tasks coming in, which can happen because of:BugsSmaller feature requestsDesign updatesand many other cases.Thus, you need a strong methodology for working through all of these tasks, verifying they’re solved in a correct manner, and marking them done. Some smaller tasks can be done fully autonomously by coding agents. However, I also found that a lot of similar-looking tasks are a bit ambiguous. If you simply ask a coding agent to fix such a task without any more input, you might find that the coding agent did not actually solve the problem, or in many cases, even worse, that the coding agent did something it wasn’t supposed to do and changed a part of your application that you didn’t intend to change.Due to the challenges I mentioned here:Solving a lot of smaller tasksAmbiguities in smaller tasksYou need a good methodology for completing all these tasks, which is what I’ll cover in the following sections.My coding methodologyNow I’ll cover my coding methodology to more effectively solve a lot of these tasks. I’ll take you through my high-level pipeline and the philosophy and mindset that I have for solving these tasks.The pipeline looks as follows:The issue is posted, typically through SlackAn agent picks up the task and creates a Linear issue for itI have a Claude Code session to deal with all of these tasks, typically for a specific time period, for one specific day in my case.I triage the tasks through my Claude Code session. If it’s a simple quick fix, I use the Claude Code session to fix it. If it’s a bigger issue, I have the agent in the session create a hand-off up front and work on a completely separate thread to solve the issue because it needs more human interaction.I make Claude Code create an HTML report of all the smaller issues that we want to work through. If it has any questions for me, I need to clarify them and give it guidelines on how to complete the tasks.Claude Code spins up a sub-agent for each smaller task and drives it to devOnce it’s in dev, I receive another HTML report on how to test the feature, and I check if it was implemented correctly. If this is the case, the task is marked done. If not, I iterate until it’s implemented correctlyIssue triagingFirst, now I want to talk about the first five steps in my pipeline, which can be summarized into issue triaging. So, basically, you should, of course, have a common place where all the feedback is posted. Slack is a great channel to do that, but you can also use any other messaging app, of course. I then have an automatic bot that creates Linear issues or tickets. I use Linear because it’s a good and clean interface for interacting with coding agents. They have automatic updates on task progress, and you can easily post updates to any task and keep a good overview of the projects you’re working on.Once a Linear ticket or issue has been created, they are now accessible to my coding agent. I typically start one Claude Code session per day for smaller tasks. So I have an August 15th session, an August 16th session, and so on. But of course, you can adapt this to any time period that you prefer.Once I’m in the Claude Code session, I ask it to read through the Linear tickets from that day or Slack and find all the tasks and map them out to an HTML report. It should then look into each task and present me with the report, with details about the task, which I read through. I give Claude Code any input that it should have to complete a certain task; for example, I try to clarify any design decisions or how something should be implemented. Also, if it’s a bigger task, which sometimes comes in, then I ask Claude to make a handoff, because I wanna do bigger tasks in a separate thread.The reason I want to do bigger tasks in a separate thread is that they require more human input, and when they require this, it gets very messy if I have it in the main Claude Code sessions where I do all of the smaller tasks. It’s better to have it in a separate session where all the questions the coding agent has for me are centralized in one location, and I can interact with the coding agent there. I simply find that it’s a more efficient way to complete bigger tasks.After this, I’m done with the issue triaging.Effectively solving the tasksNow let’s talk about point number 6, which is about how I effectively solve all of these smaller tasks. The simple way I do it is that I ask Claude Code explicitly to spin up sub-agents to complete each task individually. When you do this, it’s very important that you instruct Claude Code to spin up sub-agents in separate worktrees so that the sub-agents don’t interfere with each other. And this is a great way to do it because Claude spins up one sub-agent per task that you’re working on, and it’s very easy to keep an overview of all the sub-agents. You can basically see them in the menu in the CLI. If you want to dive into one specific sub-agent, which admittedly is something I do quite rarely, you can also just click on it and see what’s going on there.Then I basically let Claude Code continue working on each subtask, asking it, of course, to implement it correctly, verify its own work, run a code review, and drive it to dev immediately. In most cases, I ask Claude Code to simply drive it directly to dev. Though, if it’s a task such as a design task where I know agents can make mistakes, I might have the sub-agent spin up a localhost server and verify the work there before I ask the model to drive it to dev.Verifying the workThe last step is, of course, to verify the work. I find that in most cases, it’s worth just spending 30 seconds to 1 minute verifying the work for one task. In most cases, Claude has implemented it correctly, but I do find that it’s very hard to know which tasks are likely to be implemented incorrectly, and I thus do spend the time verifying the work manually.However, I have optimized the way I verify the work. To verify the work, I basically ask Claude Code to present me with an HTML report with each task that it implemented and exactly how I can test the task. This should include the original Slack message or Linear issue quoted verbatim. It should include a link to the exact page where I can test the issue. For example, if you wanted to fix the design in the chatbot functionality, the AI should give you the link to a specific chatbot thread, so you can check it out there and you don’t have to navigate the product yourself.I can basically then just go through the checklist that the agent has provided me in the HTML report and verify the work very easily. If I deem the work to be implemented correctly, I say that the task is verified and it can be set to done because it’s already in dev most of the time. If it’s not, I give the agent feedback on what it did incorrectly, ask it to implement it, and come back to me with a new HTML report once it’s fixed so I can test it again.ConclusionThis is basically my problem-solving pipeline for coding efficiently with Claude Code. I think all the steps that are covered in this article are very important, as they each contribute to the next step being completed efficiently. For example, issue triaging is a very important prerequisite for a single Claude Code session to be able to spin up sub-agents to complete all of the smaller issues. And then having the sub-agents is, of course, very important, and having an effective way of verifying the work with HTML reports is critical to keep testing speed up with implementation speed. I hope you learned something from this article and try implementing some of this problem-solving pipeline into your own programming workflows, as I do believe this can be a very effective way of increasing speed when developing products.👋 Get in Touch👉 My free eBook and Webinar:🚀 10x Your Engineering with LLMs (Free 3-Day Email Course)📚 Get my free Vision Language Models ebook💻 My webinar on Vision Language Models👉 Find me on socials:💌 Substack🔗 LinkedIn🐦 X / Twitter

Read More »

AI models flub these intelligence tests. Can you fare any better?

Puzzles and games have been central to AI development since the very beginning. Just as we humans like to test our smarts with crosswords or logic puzzles, developers can test how far models have advanced with a gaming gauntlet. The term “machine learning” was popularized in a 1959 article by the IBM computer scientist Arthur Samuel about an algorithm that learned to play checkers. Chess and the Chinese board game Go are famous AI test beds too.  Judged purely on its puzzling skills, AI is improving a lot—and quickly. In late 2024, a team of scientists from Columbia University showed that even the best models could figure out only 18% of the infamous New York Times Connections puzzles; by early 2025, some models could solve them near perfectly every time.  But puzzles do more than just highlight the inexorable advance of AI capabilities. Seeing where models succeed and fail—and where we humans still beat them—can provide a useful window into the technology’s strengths and weaknesses. Despite advances, today’s models still fumble: Subtle changes in classic riddles often trip them up, and visual puzzles are a particular weak spot.  Here you’ll have the chance to test your wits on puzzles that have stumped models at one time or another. Some might be as tricky for you as they were for the AI; others are so simple that they’ll have you doubting whether AI is really intelligent at all. Each one highlights at least one way in which machine and human cognition differ. If you ace the test, you’ll have proved that you can out-puzzle an AI—at least for now.  Spatial Reasoning Let’s start with a domain where humans have a huge advantage: spatial reasoning. If you’ve ever taken an IQ test, you may have done a mental rotation problem. These puzzles ask you to determine whether different images represent the same objects from different angles. Though today’s language models typically have the ability to analyze visual inputs, they still fail abysmally at these puzzles. For all the talk of how world models can help AI understand physical environments, LLMs still don’t seem to be able to manipulate 3D objects the way spatial thinkers like architects and mechanical engineers can. Mental Rotation Instructions: Choose the answer that shows the object in the prompt, but from a different angle. In each case, there’s only one correct answer!

.cst-large,
.cst-default {
width: 100%;
}

@media (max-width: 767px) {
.cst-block {
overflow-x: hidden;
}
}

@media (min-width: 630px) {
.cst-large {
margin-left: -25%;
width: 150%;
}

@media (min-width: 960px) {
.cst-large {
margin-left: -16.666666666666664%;
width: 140.26%;
}
}

@media (min-width: 1312px) {
.cst-large {
width: 145.13%;
}
}
} Memory & Adaptability Frontier LLMs have extraordinary memories; they were exposed to a monstrous volume of facts during training and can recite many of them faithfully. That’s an asset for outcompeting humans at trivia, but it can also be a liability. When a puzzle closely resembles one a model saw during training, the model may whiz by key differences and respond with what it memorized.  This held true in a 2024 study in which researchers from Google and the University of Illinois Urbana-Champaign trained and tested models on slight variations of a classic type of puzzle called Knights and Knaves. In these problems, some characters always tell the truth and others always lie, and you have to figure out who’s who. The same principle may be at work in a test called SimpleBench. These questions resemble more complicated problems that models likely encountered in training. Humans spot the trick, but even top-tier models trip.
Knights and Knaves Instructions: The only thing you need to know to solve these puzzles is that knights always tell the truth and knaves always lie. Determine who’s what on the basis of what each character says.
.cst-large,
.cst-default {
width: 100%;
}

@media (max-width: 767px) {
.cst-block {
overflow-x: hidden;
}
}

@media (min-width: 630px) {
.cst-large {
margin-left: -25%;
width: 150%;
}

@media (min-width: 960px) {
.cst-large {
margin-left: -16.666666666666664%;
width: 140.26%;
}
}

@media (min-width: 1312px) {
.cst-large {
width: 145.13%;
}
}
}

1

You have met a group of two islanders.
Their names are Edward and Wallace.

Wallace says:
Edward tells the truth.

Edward says:
Wallace and I are the same type.

2

You have met a group of three islanders.
Their names are Joseph, Francine, and Alice.

Francine says:
Joseph is a knave.

Francine says:
Alice tells the truth.

Alice says:
Joseph is not my type.

3

You have met a group of three islanders.
Their names are Robert, Vincent, and Michelle.

Michelle says:
Robert always lies.

Vincent says:
Michelle is truthful.

Robert says:
Vincent is untruthful.

Robert says:
Vincent is not my type.

SimpleBench Instructions: Read these SimpleBench problems carefully, and you should be able to figure out the answers in no time.
.cst-large,
.cst-default {
width: 100%;
}

@media (max-width: 767px) {
.cst-block {
overflow-x: hidden;
}
}

@media (min-width: 630px) {
.cst-large {
margin-left: -25%;
width: 150%;
}

@media (min-width: 960px) {
.cst-large {
margin-left: -16.666666666666664%;
width: 140.26%;
}
}

@media (min-width: 1312px) {
.cst-large {
width: 145.13%;
}
}
}

1

Beth places four
whole ice cubes in a
frying pan at the start of the
first minute, then five at the
start of the second minute
and some more at the start
of the third minute, but none
in the fourth minute. If the
average number of ice cubes
per minute placed in the pan
while it was frying a crispy
egg was five, how many
whole ice cubes can be
found in the pan at the end
of the third minute?

A
30

B
0

C
20

D
10

E
11

F
5

2

A juggler throws a
solid blue ball a meter
in the air and then a solid
purple ball (of the same size)
two meters in the air. She
then climbs to the top of a
tall ladder carefully, balancing a yellow balloon on her
head. Where is the purple
ball most likely now, in relation to the blue ball?

A
At the same height as the blue ball

B
At the same height as the yellow balloon

C
Inside the blue ball

D
Above the yellow balloon

E
Below the blue ball

F
Above the blue ball

Abstract & Visual Reasoning AI doesn’t just bungle visual problems in 3D—two dimensions can trip it up as well. That’s a major factor in how well models do on the most famous ­puzzle-based benchmark, ARC-AGI. These problems require you to infer abstract, general rules from a set of examples. Models do better on ARC puzzles when they receive each grid not as an image but as a string of numbers that encodes the color of each cell.  Research suggests that even when models answer ARC-AGI questions correctly, they often do so using byzantine and non-­generalizable rules, whereas humans draw on simple visual concepts. Despite these disadvantages, models have gotten quite good at ARC-AGI over the past year, but some puzzles—such as the one printed here—still stump them. ARC-AGI Instructions: Study the three pairs of grids shown below to figure out the rule that dictates how the ones on the left transform into the ones on the right. Then get out your markers or colored pencils and fill in the fourth grid using that rule. (The solution is the same no matter which way the grids are oriented.)
.cst-large,
.cst-default {
width: 100%;
}

@media (max-width: 767px) {
.cst-block {
overflow-x: hidden;
}
}

@media (min-width: 630px) {
.cst-large {
margin-left: -25%;
width: 150%;
}

@media (min-width: 960px) {
.cst-large {
margin-left: -16.666666666666664%;
width: 140.26%;
}
}

@media (min-width: 1312px) {
.cst-large {
width: 145.13%;
}
}
}

Now you try it

Your answer

Intuition It’s not just AI models that fall into traps. We humans have our own cognitive foibles, many of which AI does not share. Psychologists have designed problem suites that invert the SimpleBench phenomenon: For these questions, humans often give knee-jerk answers, whereas models will respond deliberatively. Some of the problems exploit errors in the ways that we intuitively do math; others are phrased so as to suggest obvious answers that fall apart if the question is read carefully. 
Lightning Round Instructions: Answer the questions below as quickly as you can.
.cst-large,
.cst-default {
width: 100%;
}

@media (max-width: 767px) {
.cst-block {
overflow-x: hidden;
}
}

@media (min-width: 630px) {
.cst-large {
margin-left: -25%;
width: 150%;
}

@media (min-width: 960px) {
.cst-large {
margin-left: -16.666666666666664%;
width: 140.26%;
}
}

@media (min-width: 1312px) {
.cst-large {
width: 145.13%;
}
}
}

1

In a cave, there is a colony of bats whose population doubles each day. Given that it takes 60 days for the entire cave to be filled with bats, how many days would it take for the cave to be half-filled with bats?

2

In what famous novel does Alice state “I’m late, I’m late, for a very important date”?

Increasing Complexity In some cases, whether an LLM can complete a puzzle is a matter of scale. One study from researchers at Apple found that LLMs can ace simple versions of the Tower of Hanoi problem, which involves moving a stack of disks one at a time without ever putting a larger disk atop a smaller one, and river-crossing puzzles, in which a group of people must traverse a river according to certain rules. But only up to a point: As the number of disks or people hits six and higher, the models began to falter. In another study, researchers at the University of Washington, Stanford University, and the Allen Institute for AI observed that LLMs struggle similarly with logic grid puzzles, which require deducing the attributes of a set of individuals from a list of clues. The Apple paper went viral, but commentators questioned whether the results reveal a unique limitation of LLM reasoning—or just that it’s normal to make errors as complexity piles up. The River Instructions: Using the scenario provided, plan the trips necessary to get everyone across the river. 

.cst-large,
.cst-default {
width: 100%;
}

@media (max-width: 767px) {
.cst-block {
overflow-x: hidden;
}
}

@media (min-width: 630px) {
.cst-large {
margin-left: -25%;
width: 150%;
}

@media (min-width: 960px) {
.cst-large {
margin-left: -16.666666666666664%;
width: 140.26%;
}
}

@media (min-width: 1312px) {
.cst-large {
width: 145.13%;
}
}
}

Three FBI agents and their three informants need to cross a river. They have a rowboat that can fit only two people, though it can be rowed by only one. Each agent will refuse to leave their informant on the same bank as other agents without them present—even if the informant never steps out of the boat and onto the bank. How can all six make it across?

Logic Grid Instructions: Using the list of clues, determine who lives in each house and what style of music each person enjoys. There is only one possible solution. You may find it helpful to fill out the grid below to keep track of your deductions.
.cst-large,
.cst-default {
width: 100%;
}

@media (max-width: 767px) {
.cst-block {
overflow-x: hidden;
}
}

@media (min-width: 630px) {
.cst-large {
margin-left: -25%;
width: 150%;
}

@media (min-width: 960px) {
.cst-large {
margin-left: -16.666666666666664%;
width: 140.26%;
}
}

@media (min-width: 1312px) {
.cst-large {
width: 145.13%;
}
}
}

The Neighborhood

There are 4 houses, numbered 1 to 4 from left to right, as seen from across the street.
Each house is occupied by a different person: Peter, Eric, Arnold, or Alice.
Each resident has a favorite type of music: jazz, rock, classical, or pop.

Alice is directly left of Peter.
The person who loves classical music is directly left of Peter.
Arnold loves jazz music.
The person who loves rock music is not in the second house.
The person who loves rock music is directly left of the person who loves pop music.

Click a cell to mark an X, click again for a check mark.

Grace Huckins is an AI reporter at MIT Technology Review. They have a PhD in neuroscience. Credits: Mental Rotation: CC BY 4.0. Stogiannidis, Ilias, Steven McDonagh, Sotirios A. Tsaftaris. Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models (copyright 2025); illustrations by John MacNeill. Knights & knaves: Courtesy Dan MacKinnon. Simplebench: CC BY 4.0. SimpleBench Team. The Text Benchmark in which Unspecialized Human Performance Exceeds that of Current Frontier Models (copyright 2024). ARC-AGI: Courtesy ARC Prize Foundation. Lightning round: CC BY 4.0. Hagendorff, Thilo, Sarah Fabi, Michal Kosinski. Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in ChatGPT. Nat Comput Sci 3, 833–838 (copyright 2023). The river: Adapted from Propositiones ad Acuendos Juvenes, Alcuin of York (ca. 800 CE). Logic grid: Apache License 2.0. Lin, Bill Y., Ronan Le Bras, Kyle Richardson, et al. ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning (copyright 2025)

Read More »

Raised on AI

When my oldest child was born, I immediately set up Gmail and Twitter accounts in her name. I broadly announced her birth online and proceeded to plaster her photo across all sorts of platforms. In short, I began creating her digital footprint long before she could stand on her own two feet.  Fast-forward a couple of years to when my second kid came, and I had essentially the opposite reaction. I wanted to make sure I preserved her privacy. I didn’t want her birthday to be a matter of public record or her face to feed the algorithms. In time, I would go back and scrub much of the early footprint I had created for my first child as well.  What happened? I watched the promise of the early internet give way to the reality of its potential for abuse. I myself was already fully in the throes of smartphone and social media obsession. My wife, a pediatric nurse, grew increasingly alarmed at the number of children admitted to her hospital struggling with the effects of things like body dysmorphia or cyberbullying as a result of interactions on social media. We became those parents. The ones whose kids carry flip phones and aren’t on TikTok.  We are not alone in this. A surprising—maybe troubling—number of people I know who work at big tech companies also keep their kids at arm’s distance from technology. They lock down their phones, if they have phones at all, and keep them off social media. Hell, even Mark Zuckerberg doesn’t publicly post his children’s faces on Facebook or Instagram.  
If the desire to limit kids’ use of technology was once a subcurrent, it has become a raging flood. Jonathan Haidt’s best-selling 2024 book The Anxious Generation helped propel the issue into the mainstream (despite criticisms of his conclusions from some developmental psychologists). Last year, Australia became the first country to enact a social media ban for children under 16. Other nations, from Austria to Indonesia, have followed suit, announcing similar bans. The US Supreme Court recently upheld an age verification law in Texas, which acts as a de facto ban, and several states have flirted with their own measures. School districts all over the country are banning educational devices like iPads and Chromebooks in favor of actual books. Kids themselves seem to be embracing this tech skepticism too: The hottest gadget among the Gen Alpha set is a vintage Sony Walkman. We have to prepare our children to live in the actual world we have actually created, not the one we wish we had. Yet there is no hiding from technology. It permeates nearly everything, everywhere. And so we have to prepare our children to live in the actual world we have actually created, not the one we wish we had. How can we help kids survive and thrive in what we have wrought? 
It’s a question that feels all the more urgent in the era of AI. To help answer it, we brought in the editors of Anyway—an utterly fantastic magazine for teens and tweens that is so good in large part because it meets them where they are. (If there is a teen in your life, I highly recommend it.) They helped us with the stories you’ll see in this issue, and they asked kids to share in their own words how they are feeling about AI and what’s to come. What those young people told Anyway was complex, fascinating, and, to an incredible extent, thoughtful and sophisticated.  Meanwhile, I’ve loosened the digital tether—just a bit. I still don’t post many photos of my kids online, and I remain abundantly concerned about the perils of social media and AI.  But when my older daughter started high school, we retired the flip phone in favor of an iPhone. And my younger one now sports an Apple Watch. These devices have opened up the world to them in all sorts of ways. They help forge new friendships, building relationships that move seamlessly between digital and physical spaces. They allow my kids to roam free—or at least more freely—beyond the known spaces of our neighborhood and across the city.  Along the way, they’re learning and testing boundaries, just the way they’re supposed to. I guess you could say I am too. 

Read More »

Hugging Face hack could indicate cultural issues at OpenAI

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. By now you’ve probably heard about last month’s major AI security incident, in which OpenAI agents escaped their sandbox and hacked into the AI platform Hugging Face while trying to cheat on a test. It’s a wild story. On Wednesday, OpenAI released a postmortem technical report on the incident, which I wrote about here.  The day before OpenAI released that report, I spoke with David Krueger, a computer science professor and prominent alignment expert who took leave from the University of Montreal to found and lead an AI safety nonprofit called Evitable. He said what he had really hoped to see in the report was an analysis of the human factors behind the incident. “When you look at accidents and incidents, oftentimes people try to find the technical source of failure, but that can give a very inaccurate and misleading sense of why the failure occurred,” he said. “If people are just cutting corners all the time, if people are not in a culture that prioritizes safety and has appropriate incentives and structures, [accidents] are kind of bound to happen.”
The report did not meet Krueger’s hopes. Its 38 pages detail a multi-month progression of agent misbehavior that culminated in the Hugging Face hack, explore the technical reasons why that misbehavior occurred, and enumerate the steps being taken to prevent similar events in the future. But there’s no consideration of the role that company culture may have played in the incident, and the report includes few references to specific human errors.  That’s all the more concerning because the references to human error in the report suggest that significant cultural issues could be at play. Back in May, models in training figured out how to communicate with one another via an improvised message board, and an OpenAI team observed the behavior. Because that behavior occurred during training, the models learned that secret interagent communication was a viable strategy for completing tasks—but rather than restarting the training process, the team allowed the models to move forward with that risky information encoded in their weights.
When those models were tested in late June, they again created a message board, which enabled the Hugging Face attack. This message board, too, was discovered, but the employees who responded determined that evaluation could continue, and the report suggests that no one higher up the chain of command realized what was going on until it was far too late. “For this to have gotten this out of control in this way requires a very long series of failures, a cascading set of failures that cause an increasingly large footprint that if at any point a human notices and raises the alarm, this should end,” says Zvi Mowshowitz, a popular AI safety writer on Substack who has drawn attention to OpenAI’s failure to halt training after the first message board was discovered. According to the report, OpenAI employees noticed what was happening at multiple points—and either failed to raise the alarm or were not heard when they did. What OpenAI’s report fails to address is why a company that develops such high-risk systems did not prevent this severe communication breakdown, though Mowshowitz has his suspicions. “All these different failures are all pointing in the same direction, which is that the safety culture at OpenAI doesn’t exist or is anemically weak,” he says. Of course, just because we don’t see a deep analysis of safety factors in the report doesn’t mean that OpenAI isn’t conducting one internally. But in an email to MIT Technology Review, Johns Hopkins University professor emeritus and organizational safety expert Kathleen Sutcliffe expressed concern that the public report did not include any reflection on the company’s practices and culture. “The ways in which people interact—the daily habits, routines, and practices we engage in in our organizational lives—affect our abilities to be alert and aware of unfolding events, our abilities to make sense of what we see, and ultimately our abilities to cope with events as they unfold,” she wrote.  In response to questions about whether and how the company is reflecting on its safety culture, OpenAI referred MIT Technology Review back to the technical report.  We do know that at least some high-level reflection on safety procedures has taken place at OpenAI, because the technical report does make clear that the company is updating its protocols for responding to safety incidents. But culture change is a tricky problem, and without more information from the company, it’s difficult to say whether strengthened response protocols alone will do much to prevent a future crisis. In its report, OpenAI spends a great deal of time reflecting on the failures in alignment between the AI models the company trains and tests and the humans who run them. But even bigger alignment problems may exist in the disconnect between company culture and the public interest. And as tough as technical AI research might be, fixing those problems could prove far harder.

Read More »

Nvidia unveils alternative high bandwidth technology to bolster AI cards

Nvidia is once again making its own parts for AI rather than relying on the rest of the industry to do it. In this case, it has introduced a customized High Bandwidth Memory architecture, NVHBM, that the company says can deliver significantly more bandwidth, lower power consumption and offer more usable silicon area than conventional HBM4E. To be sure, Nvidia will not be making the memory. It doesn’t make its own chips and has no foundry, after all. One of the big three memory makers – Micron, SK Hynix, or Samsung — we’ll actually make the chips. Nvidia is just designing them. The new memory was announced in a blog post on the same day as the company’s quarterly earnings call. It is positioned as an expansion of Nvidia’s NVLink Fusion platform. NVHBM is aimed directly at hyperscalers and AI companies developing their own custom accelerators, or XPUs. Amazon’s Annapurna Labs will be the first announced partner to work with Nvidia on the technology.

Read More »

Private AI cloud, agentic infrastructure dominate VMware Explore

In fact, Flexential is now building its own set of offerings on top of the VMware AI Factory, set to be released early next year. “Basically, you can do everything there,” Cook tells Network World. “You can bring your own model. You can spin up your agents. You can orchestrate it. And they’ve got observability, which is critical.” Flexential will bundle that with hosting and services for a complete private AI solution. “Their stuff sits next to most of the data in the enterprise today, and that latency, that proximity, is important,” he says. Running AI on premises instead of in the cloud can offer cost, latency, and compliance benefits, but it can also be a management challenge, and organizations are still struggling to find the right balance. According to an Omdia survey of 1,201 IT leaders released in August, 96% currently use a mix of cloud, on-premise, and edge infrastructure for AI. As share of workload, 59% of inferencing workloads currently run in the cloud, and 41% on-prem. In three years, companies expect to run 63% of AI inferencing workloads in the cloud and 37% on-prem. Meanwhile, other workloads are moving in the opposite direction—according to the survey, 60% of companies have repatriated some workloads back from the cloud to on-prem. For companies running AI workloads on prem, integrating AI security with a virtualization platform makes sense, says Ryan Sheehan, senior vice president of advanced solutions at SHI International, an IT consultancy. “That’s where I think VMware is in a very good position,” he tells Network World. “They’re close to the infrastructure.” That puts VMware in the right place to enforce controls. “They’re at the hypervisor,” he says. “They are the hypervisor.”

Read More »

Tailscale expands from VPN into a full connectivity platform

DNS Filtering by Control D packages an existing integration into a single purchase. Control D is a DNS filtering service that blocks malicious, phishing, and unapproved domains before a device connects to them. The new add-on lets customers apply Control D’s filtering profiles directly through Tailscale’s own policy engine, by user, group, tag, or device, rather than managing two separate consoles. Tailscale PAM addresses a different piece of that problem: privileged access to specific infrastructure rather than DNS destinations. Tailscale acquired Border0 in March 2026, the technology behind Tailscale PAM. Tailscale PAM lets teams grant one-click access to specific servers, databases, Kubernetes clusters, and web applications, without handing out standing passwords or API keys. Every session, human or AI agent, gets logged and can be scoped to a specific time window for audits and compliance. Aperture turns VPN identity into an AI gateway Aperture is Tailscale’s AI gateway. It gives an AI agent an identity on a tailnet, the same way a device or a person already gets one, then routes that agent’s model calls and tool use through Tailscale’s private network instead of the open internet. That extends the same core idea behind Tailscale’s original VPN, identity-based access instead of network-based access, to AI agents specifically.

Read More »

CBRE: Record Data Center Construction Fails to Ease Capacity Crunch

The North American data center industry built at record scale during the first half of 2026. It still wasn’t enough to loosen the market. Primary-market supply increased 33.7% year over year to a record 10,903 MW, according to CBRE’s newly released North America Data Center Trends H1 2026. Yet vacancy moved in the opposite direction, falling from 1.6% a year earlier to another record low of 1.4%. Net absorption reached 1,456.2 MW, up 11.7%, as hyperscale cloud and AI infrastructure operators continued competing for increasingly scarce blocks of contiguous power and capacity. Construction climbed 24.8% to a record 7,481.1 MW, surpassing the previous peak of 6,350.1 MW set in the second half of 2024. Perhaps the most telling number in the report is 80.4%. That is the share of primary-market capacity under construction that has already been preleased, up from 74.3% a year ago. CBRE estimates that less than 1,500 MW of all capacity currently under construction remains available—roughly six months of demand at the current absorption rate. The resulting picture is less one of a construction shortage than a race between infrastructure delivery and an AI demand curve that keeps absorbing capacity before it reaches the market. And increasingly, CBRE argues, securing the megawatts is only part of that race. Power, Permitting and Local Approval Converge Power availability and infrastructure delivery timelines remain the primary determinants of where data centers can be built. But CBRE’s H1 report places another constraint alongside them: community acceptance. “Local opposition has also become a serious obstacle,” the report states, with community resistance, zoning disputes and entitlement delays increasingly capable of stopping projects even after developers identify viable power and fiber. CBRE goes further in its outlook, describing community engagement as a development constraint now “on par with power procurement.” In practical site-selection terms,

Read More »

Corvex Tests a Faster Path to Liquid-Cooled AI Infrastructure

The customer agreement expanded an earlier commitment and includes dedicated high-speed storage and CPUs in addition to GPUs. Corvex initially delivered capacity during the first quarter and continued deployment through the second and third quarters. By its Aug. 14 earnings update, the company said the multi-year Blackwell agreement had been fully delivered. Corvex reported approximately $22 million in contracted annualized recurring revenue from compute that was live and accepted by customers. The expansion was financed through debt, customer prepayments and cash on hand rather than additional equity issuance. But the most noteworthy aspect of the project may be the deployment itself. Corvex installed high-density, liquid-cooled NVIDIA HGX B200 systems inside an existing air-cooled data center and says it commissioned the capacity approximately two weeks after the equipment arrived. Instead of rebuilding the facility around a central liquid-cooling system, Corvex worked with Lenovo to use Lenovo Neptune liquid-to-air cooling technology. Liquid removes heat from the servers and transfers it to the existing air-cooled facility infrastructure. The cluster uses Lenovo ThinkSystem systems equipped with NVIDIA HGX B200 GPUs, NVIDIA Quantum-2 QM9700 InfiniBand for GPU traffic, NVIDIA Spectrum SN5600 switches for storage networking and SN2201 switches for management traffic. Corvex says the design allowed it to place high-density Blackwell infrastructure into the existing facility without a conventional facility-wide liquid-cooling conversion. Lenovo, in a case study of the deployment, contrasts the approximately two-week commissioning period with what it describes as typical data center upgrade timelines of seven to 12 months or more. That has significance beyond a single cluster. Power availability and suitable data center capacity increasingly constrain GPU deployment. If the approach proves repeatable, liquid-to-air cooling could allow some existing air-cooled facilities with sufficient power and other supporting infrastructure to accommodate higher-density AI systems without first undergoing a full central liquid-cooling conversion. For

Read More »

Stay Ahead with the Paperboy Newsletter

Your weekly dose of insights into AI, Bitcoin mining, Datacenters and Energy indusrty news. Spend 3-5 minutes and catch-up on 1 week of news.

Smarter with ONMINE

Streamline Your Growth with ONMINE