Stay Ahead, Stay ONMINE

LLM + RAG: Creating an AI-Powered File Reader Assistant

Introduction AI is everywhere.  It is hard not to interact at least once a day with a Large Language Model (LLM). The chatbots are here to stay. They’re in your apps, they help you write better, they compose emails, they read emails…well, they do a lot. And I don’t think that that is bad. In fact, my opinion is the other way – at least so far. I defend and advocate for the use of AI in our daily lives because, let’s agree, it makes everything much easier. I don’t have to spend time double-reading a document to find punctuation problems or type. AI does that for me. I don’t waste time writing that follow-up email every single Monday. AI does that for me. I don’t need to read a huge and boring contract when I have an AI to summarize the main takeaways and action points to me! These are only some of AI’s great uses. If you’d like to know more use cases of LLMs to make our lives easier, I wrote a whole book about them. Now, thinking as a data scientist and looking at the technical side, not everything is that bright and shiny.  LLMs are great for several general use cases that apply to anyone or any company. For example, coding, summarizing, or answering questions about general content created until the training cutoff date. However, when it comes to specific business applications, for a single purpose, or something new that didn’t make the cutoff date, that is when the models won’t be that useful if used out-of-the-box – meaning, they will not know the answer. Thus, it will need adjustments. Training an LLM model can take months and millions of dollars. What is even worse is that if we don’t adjust and tune the model to our purpose, there will be unsatisfactory results or hallucinations (when the model’s response doesn’t make sense given our query). So what is the solution, then? Spending a lot of money retraining the model to include our data? Not really. That’s when the Retrieval-Augmented Generation (RAG) becomes useful. RAG is a framework that combines getting information from an external knowledge base with large language models (LLMs). It helps AI models produce more accurate and relevant responses. Let’s learn more about RAG next. What is RAG? Let me tell you a story to illustrate the concept. I love movies. For some time in the past, I knew which movies were competing for the best movie category at the Oscars or the best actors and actresses. And I would certainly know which ones got the statue for that year. But now I am all rusty on that subject. If you asked me who was competing, I would not know. And even if I tried to answer you, I would give you a weak response.  So, to provide you with a quality response, I will do what everybody else does: search for the information online, obtain it, and then give it to you. What I just did is the same idea as the RAG: I obtained data from an external database to give you an answer. When we enhance the LLM with a content store where it can go and retrieve data to augment (increase) its knowledge base, that is the RAG framework in action. RAG is like creating a content store where the model can enhance its knowledge and respond more accurately. User prompt about Content C. LLM retrieves external content to aggregate to the answer. Image by the author. Summarizing: Uses search algorithms to query external data sources, such as databases, knowledge bases, and web pages. Pre-processes the retrieved information. Incorporates the pre-processed information into the LLM. Why use RAG? Now that we know what the RAG framework is let’s understand why we should be using it. Here are some of the benefits: Enhances factual accuracy by referencing real data. RAG can help LLMs process and consolidate knowledge to create more relevant answers  RAG can help LLMs access additional knowledge bases, such as internal organizational data  RAG can help LLMs create more accurate domain-specific content  RAG can help reduce knowledge gaps and AI hallucination As previously explained, I like to say that with the RAG framework, we are giving an internal search engine for the content we want it to add to the knowledge base. Well. All of that is very interesting. But let’s see an application of RAG. We will learn how to create an AI-powered PDF Reader Assistant. Project This is an application that allows users to upload a PDF document and ask questions about its content using AI-powered natural language processing (NLP) tools.  The app uses Streamlit as the front end. Langchain, OpenAI’s GPT-4 model, and FAISS (Facebook AI Similarity Search) for document retrieval and question answering in the backend. Let’s break down the steps for better understanding: Loading a PDF file and splitting it into chunks of text. This makes the data optimized for retrieval Present the chunks to an embedding tool. Embeddings are numerical vector representations of data used to capture relationships, similarities, and meanings in a way that machines can understand. They are widely used in Natural Language Processing (NLP), recommender systems, and search engines. Next, we put those chunks of text and embeddings in the same DB for retrieval. Finally, we make it available to the LLM. Data preparation Preparing a content store for the LLM will take some steps, as we just saw. So, let’s start by creating a function that can load a file and split it into text chunks for efficient retrieval. # Imports from langchain_community.document_loaders import PyPDFLoader from langchain.text_splitter import RecursiveCharacterTextSplitter def load_document(pdf): # Load a PDF “”” Load a PDF and split it into chunks for efficient retrieval. :param pdf: PDF file to load :return: List of chunks of text “”” loader = PyPDFLoader(pdf) docs = loader.load() # Instantiate Text Splitter with Chunk Size of 500 words and Overlap of 100 words so that context is not lost text_splitter = RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=100) # Split into chunks for efficient retrieval chunks = text_splitter.split_documents(docs) # Return return chunks Next, we will start building our Streamlit app, and we’ll use that function in the next script. Web application We will begin importing the necessary modules in Python. Most of those will come from the langchain packages. FAISS is used for document retrieval; OpenAIEmbeddings transforms the text chunks into numerical scores for better similarity calculation by the LLM; ChatOpenAI is what enables us to interact with the OpenAI API; create_retrieval_chain is what actually the RAG does, retrieving and augmenting the LLM with that data; create_stuff_documents_chain glues the model and the ChatPromptTemplate. Note: You will need to generate an OpenAI Key to be able to run this script. If it’s the first time you’re creating your account, you get some free credits. But if you have it for some time, it is possible that you will have to add 5 dollars in credits to be able to access OpenAI’s API. An option is using Hugging Face’s Embedding.  # Imports from langchain_community.vectorstores import FAISS from langchain_openai import OpenAIEmbeddings from langchain.chains import create_retrieval_chain from langchain_openai import ChatOpenAI from langchain.chains.combine_documents import create_stuff_documents_chain from langchain_core.prompts import ChatPromptTemplate from scripts.secret import OPENAI_KEY from scripts.document_loader import load_document import streamlit as st This first code snippet will create the App title, create a box for file upload, and prepare the file to be added to the load_document() function. # Create a Streamlit app st.title(“AI-Powered Document Q&A”) # Load document to streamlit uploaded_file = st.file_uploader(“Upload a PDF file”, type=”pdf”) # If a file is uploaded, create the TextSplitter and vector database if uploaded_file :     # Code to work around document loader from Streamlit and make it readable by langchain     temp_file = “./temp.pdf”     with open(temp_file, “wb”) as file:         file.write(uploaded_file.getvalue())         file_name = uploaded_file.name     # Load document and split it into chunks for efficient retrieval.     chunks = load_document(temp_file)     # Message user that document is being processed with time emoji     st.write(“Processing document… :watch:”) Machines understand numbers better than text, so in the end, we will have to provide the model with a database of numbers that it can compare and check for similarity when performing a query. That’s where the embeddings will be useful to create the vector_db, in this next piece of code. # Generate embeddings     # Embeddings are numerical vector representations of data, typically used to capture relationships, similarities,     # and meanings in a way that machines can understand. They are widely used in Natural Language Processing (NLP),     # recommender systems, and search engines.     embeddings = OpenAIEmbeddings(openai_api_key=OPENAI_KEY,                                   model=”text-embedding-ada-002″)     # Can also use HuggingFaceEmbeddings     # from langchain_huggingface.embeddings import HuggingFaceEmbeddings     # embeddings = HuggingFaceEmbeddings(model_name=”sentence-transformers/all-MiniLM-L6-v2″)     # Create vector database containing chunks and embeddings     vector_db = FAISS.from_documents(chunks, embeddings) Next, we create a retriever object to navigate in the vector_db. # Create a document retriever     retriever = vector_db.as_retriever()     llm = ChatOpenAI(model_name=”gpt-4o-mini”, openai_api_key=OPENAI_KEY) Then, we will create the system_prompt, which is a set of instructions to the LLM on how to answer, and we will create a prompt template, preparing it to be added to the model once we get the input from the user. # Create a system prompt     # It sets the overall context for the model.     # It influences tone, style, and focus before user interaction starts.     # Unlike user inputs, a system prompt is not visible to the end user.     system_prompt = (         “You are a helpful assistant. Use the given context to answer the question.”         “If you don’t know the answer, say you don’t know. ”         “{context}”     )     # Create a prompt Template     prompt = ChatPromptTemplate.from_messages(         [             (“system”, system_prompt),             (“human”, “{input}”),         ]     )     # Create a chain     # It creates a StuffDocumentsChain, which takes multiple documents (text data) and “stuffs” them together before passing them to the LLM for processing.     question_answer_chain = create_stuff_documents_chain(llm, prompt) Moving on, we create the core of the RAG framework, pasting together the retriever object and the prompt. This object adds relevant documents from a data source (e.g., a vector database) and makes it ready to be processed using an LLM to generate a response. # Creates the RAG      chain = create_retrieval_chain(retriever, question_answer_chain) Finally, we create the variable question for the user input. If this question box is filled with a query, we pass it to the chain, which calls the LLM to process and return the response, which will be printed on the app’s screen. # Streamlit input for question     question = st.text_input(“Ask a question about the document:”)     if question:         # Answer         response = chain.invoke({“input”: question})[‘answer’]         st.write(response) Here is a screenshot of the result. Screenshot of the final app. Image by the author. And this is a GIF for you to see the File Reader Ai Assistant in action! File Reader AI Assistant in action. Image by the author. Before you go In this project, we learned what the RAG framework is and how it helps the Llm to perform better and also perform well with specific knowledge. AI can be powered with knowledge from an instruction manual, databases from a company, some finance files, or contracts, and then become fine-tuned to respond accurately to domain-specific content queries. The knowledge base is augmented with a content store. To recap, this is how the framework works: 1️⃣ User Query → Input text is received. 2️⃣ Retrieve Relevant Documents → Searches a knowledge base (e.g., a database, vector store). 3️⃣ Augment Context → Retrieved documents are added to the input. 4️⃣ Generate Response → An LLM processes the combined input and produces an answer. GitHub repository https://github.com/gurezende/Basic-Rag About me If you liked this content and want to learn more about my work, here is my website, where you can also find all my contacts. https://gustavorsantos.me References https://cloud.google.com/use-cases/retrieval-augmented-generation https://www.ibm.com/think/topics/retrieval-augmented-generation https://python.langchain.com/docs/introduction https://www.geeksforgeeks.org/how-to-get-your-own-openai-api-key

Introduction

AI is everywhere. 

It is hard not to interact at least once a day with a Large Language Model (LLM). The chatbots are here to stay. They’re in your apps, they help you write better, they compose emails, they read emails…well, they do a lot.

And I don’t think that that is bad. In fact, my opinion is the other way – at least so far. I defend and advocate for the use of AI in our daily lives because, let’s agree, it makes everything much easier.

I don’t have to spend time double-reading a document to find punctuation problems or type. AI does that for me. I don’t waste time writing that follow-up email every single Monday. AI does that for me. I don’t need to read a huge and boring contract when I have an AI to summarize the main takeaways and action points to me!

These are only some of AI’s great uses. If you’d like to know more use cases of LLMs to make our lives easier, I wrote a whole book about them.

Now, thinking as a data scientist and looking at the technical side, not everything is that bright and shiny. 

LLMs are great for several general use cases that apply to anyone or any company. For example, coding, summarizing, or answering questions about general content created until the training cutoff date. However, when it comes to specific business applications, for a single purpose, or something new that didn’t make the cutoff date, that is when the models won’t be that useful if used out-of-the-box – meaning, they will not know the answer. Thus, it will need adjustments.

Training an LLM model can take months and millions of dollars. What is even worse is that if we don’t adjust and tune the model to our purpose, there will be unsatisfactory results or hallucinations (when the model’s response doesn’t make sense given our query).

So what is the solution, then? Spending a lot of money retraining the model to include our data?

Not really. That’s when the Retrieval-Augmented Generation (RAG) becomes useful.

RAG is a framework that combines getting information from an external knowledge base with large language models (LLMs). It helps AI models produce more accurate and relevant responses.

Let’s learn more about RAG next.

What is RAG?

Let me tell you a story to illustrate the concept.

I love movies. For some time in the past, I knew which movies were competing for the best movie category at the Oscars or the best actors and actresses. And I would certainly know which ones got the statue for that year. But now I am all rusty on that subject. If you asked me who was competing, I would not know. And even if I tried to answer you, I would give you a weak response. 

So, to provide you with a quality response, I will do what everybody else does: search for the information online, obtain it, and then give it to you. What I just did is the same idea as the RAG: I obtained data from an external database to give you an answer.

When we enhance the LLM with a content store where it can go and retrieve data to augment (increase) its knowledge base, that is the RAG framework in action.

RAG is like creating a content store where the model can enhance its knowledge and respond more accurately.

Diagram: User prompts and content using LLM + RAG
User prompt about Content C. LLM retrieves external content to aggregate to the answer. Image by the author.

Summarizing:

  1. Uses search algorithms to query external data sources, such as databases, knowledge bases, and web pages.
  2. Pre-processes the retrieved information.
  3. Incorporates the pre-processed information into the LLM.

Why use RAG?

Now that we know what the RAG framework is let’s understand why we should be using it.

Here are some of the benefits:

  • Enhances factual accuracy by referencing real data.
  • RAG can help LLMs process and consolidate knowledge to create more relevant answers 
  • RAG can help LLMs access additional knowledge bases, such as internal organizational data 
  • RAG can help LLMs create more accurate domain-specific content 
  • RAG can help reduce knowledge gaps and AI hallucination

As previously explained, I like to say that with the RAG framework, we are giving an internal search engine for the content we want it to add to the knowledge base.

Well. All of that is very interesting. But let’s see an application of RAG. We will learn how to create an AI-powered PDF Reader Assistant.

Project

This is an application that allows users to upload a PDF document and ask questions about its content using AI-powered natural language processing (NLP) tools. 

  • The app uses Streamlit as the front end.
  • Langchain, OpenAI’s GPT-4 model, and FAISS (Facebook AI Similarity Search) for document retrieval and question answering in the backend.

Let’s break down the steps for better understanding:

  1. Loading a PDF file and splitting it into chunks of text.
    1. This makes the data optimized for retrieval
  2. Present the chunks to an embedding tool.
    1. Embeddings are numerical vector representations of data used to capture relationships, similarities, and meanings in a way that machines can understand. They are widely used in Natural Language Processing (NLP), recommender systems, and search engines.
  3. Next, we put those chunks of text and embeddings in the same DB for retrieval.
  4. Finally, we make it available to the LLM.

Data preparation

Preparing a content store for the LLM will take some steps, as we just saw. So, let’s start by creating a function that can load a file and split it into text chunks for efficient retrieval.

# Imports
from  langchain_community.document_loaders import PyPDFLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter

def load_document(pdf):
    # Load a PDF
    """
    Load a PDF and split it into chunks for efficient retrieval.

    :param pdf: PDF file to load
    :return: List of chunks of text
    """

    loader = PyPDFLoader(pdf)
    docs = loader.load()

    # Instantiate Text Splitter with Chunk Size of 500 words and Overlap of 100 words so that context is not lost
    text_splitter = RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=100)
    # Split into chunks for efficient retrieval
    chunks = text_splitter.split_documents(docs)

    # Return
    return chunks

Next, we will start building our Streamlit app, and we’ll use that function in the next script.

Web application

We will begin importing the necessary modules in Python. Most of those will come from the langchain packages.

FAISS is used for document retrieval; OpenAIEmbeddings transforms the text chunks into numerical scores for better similarity calculation by the LLM; ChatOpenAI is what enables us to interact with the OpenAI API; create_retrieval_chain is what actually the RAG does, retrieving and augmenting the LLM with that data; create_stuff_documents_chain glues the model and the ChatPromptTemplate.

Note: You will need to generate an OpenAI Key to be able to run this script. If it’s the first time you’re creating your account, you get some free credits. But if you have it for some time, it is possible that you will have to add 5 dollars in credits to be able to access OpenAI’s API. An option is using Hugging Face’s Embedding. 

# Imports
from langchain_community.vectorstores import FAISS
from langchain_openai import OpenAIEmbeddings
from langchain.chains import create_retrieval_chain
from langchain_openai import ChatOpenAI
from langchain.chains.combine_documents import create_stuff_documents_chain
from langchain_core.prompts import ChatPromptTemplate
from scripts.secret import OPENAI_KEY
from scripts.document_loader import load_document
import streamlit as st

This first code snippet will create the App title, create a box for file upload, and prepare the file to be added to the load_document() function.

# Create a Streamlit app
st.title("AI-Powered Document Q&A")

# Load document to streamlit
uploaded_file = st.file_uploader("Upload a PDF file", type="pdf")

# If a file is uploaded, create the TextSplitter and vector database
if uploaded_file :

    # Code to work around document loader from Streamlit and make it readable by langchain
    temp_file = "./temp.pdf"
    with open(temp_file, "wb") as file:
        file.write(uploaded_file.getvalue())
        file_name = uploaded_file.name

    # Load document and split it into chunks for efficient retrieval.
    chunks = load_document(temp_file)

    # Message user that document is being processed with time emoji
    st.write("Processing document... :watch:")

Machines understand numbers better than text, so in the end, we will have to provide the model with a database of numbers that it can compare and check for similarity when performing a query. That’s where the embeddings will be useful to create the vector_db, in this next piece of code.

# Generate embeddings
    # Embeddings are numerical vector representations of data, typically used to capture relationships, similarities,
    # and meanings in a way that machines can understand. They are widely used in Natural Language Processing (NLP),
    # recommender systems, and search engines.
    embeddings = OpenAIEmbeddings(openai_api_key=OPENAI_KEY,
                                  model="text-embedding-ada-002")

    # Can also use HuggingFaceEmbeddings
    # from langchain_huggingface.embeddings import HuggingFaceEmbeddings
    # embeddings = HuggingFaceEmbeddings(model_name="sentence-transformers/all-MiniLM-L6-v2")

    # Create vector database containing chunks and embeddings
    vector_db = FAISS.from_documents(chunks, embeddings)

Next, we create a retriever object to navigate in the vector_db.

# Create a document retriever
    retriever = vector_db.as_retriever()
    llm = ChatOpenAI(model_name="gpt-4o-mini", openai_api_key=OPENAI_KEY)

Then, we will create the system_prompt, which is a set of instructions to the LLM on how to answer, and we will create a prompt template, preparing it to be added to the model once we get the input from the user.

# Create a system prompt
    # It sets the overall context for the model.
    # It influences tone, style, and focus before user interaction starts.
    # Unlike user inputs, a system prompt is not visible to the end user.

    system_prompt = (
        "You are a helpful assistant. Use the given context to answer the question."
        "If you don't know the answer, say you don't know. "
        "{context}"
    )

    # Create a prompt Template
    prompt = ChatPromptTemplate.from_messages(
        [
            ("system", system_prompt),
            ("human", "{input}"),
        ]
    )

    # Create a chain
    # It creates a StuffDocumentsChain, which takes multiple documents (text data) and "stuffs" them together before passing them to the LLM for processing.

    question_answer_chain = create_stuff_documents_chain(llm, prompt)

Moving on, we create the core of the RAG framework, pasting together the retriever object and the prompt. This object adds relevant documents from a data source (e.g., a vector database) and makes it ready to be processed using an LLM to generate a response.

# Creates the RAG
     chain = create_retrieval_chain(retriever, question_answer_chain)

Finally, we create the variable question for the user input. If this question box is filled with a query, we pass it to the chain, which calls the LLM to process and return the response, which will be printed on the app’s screen.

# Streamlit input for question
    question = st.text_input("Ask a question about the document:")
    if question:
        # Answer
        response = chain.invoke({"input": question})['answer']
        st.write(response)

Here is a screenshot of the result.

Screenshot of the AI-Powered Document Q&A
Screenshot of the final app. Image by the author.

And this is a GIF for you to see the File Reader Ai Assistant in action!

GIF of the File Reader AI Assistant in action
File Reader AI Assistant in action. Image by the author.

Before you go

In this project, we learned what the RAG framework is and how it helps the Llm to perform better and also perform well with specific knowledge.

AI can be powered with knowledge from an instruction manual, databases from a company, some finance files, or contracts, and then become fine-tuned to respond accurately to domain-specific content queries. The knowledge base is augmented with a content store.

To recap, this is how the framework works:

1️⃣ User Query → Input text is received.

2️⃣ Retrieve Relevant Documents → Searches a knowledge base (e.g., a database, vector store).

3️⃣ Augment Context → Retrieved documents are added to the input.

4️⃣ Generate Response → An LLM processes the combined input and produces an answer.

GitHub repository

https://github.com/gurezende/Basic-Rag

About me

If you liked this content and want to learn more about my work, here is my website, where you can also find all my contacts.

https://gustavorsantos.me

References

https://cloud.google.com/use-cases/retrieval-augmented-generation

https://www.ibm.com/think/topics/retrieval-augmented-generation

https://youtu.be/T-D1OfcDW1M?si=G0UWfH5-wZnMu0nw

https://python.langchain.com/docs/introduction

https://www.geeksforgeeks.org/how-to-get-your-own-openai-api-key

Shape
Shape
Stay Ahead

Explore More Insights

Stay ahead with more perspectives on cutting-edge power, infrastructure, energy,  bitcoin and AI solutions. Explore these articles to uncover strategies and insights shaping the future of industries.

Shape

AMD agrees to buy World Labs to fill out its AI stack

This is where World Labs fits into AMD’s ecosystem, according to Parv Sharma, Senior Research Analyst at Counterpoint Research. “World Labs builds AI that understands space, where models need to understand geometry, physics and time, unlike LLMs, which understand languages. These world models are used for training in physical AI

Read More »

Why a network digital twin is the missing piece for AI-era operations

The e-book draws an important distinction between two approaches that share the label. One emulates the network by running the actual device firmware against specific test scenarios. The other builds a deterministic mathematical model from the network’s configuration and state, computing all possible forwarding behaviors at once. The guide sums

Read More »

NetScaler admins told to patch critical zero-days in ADC and Gateway now

NetScaler appliances are an important part of many enterprise networks, providing VPN and remote access, load balancing and other application delivery services. Citrix is tracking the two exploited vulnerabilities as CVE-2026-88771 and CVE-2026-88772. It has released fixes in NetScaler ADC and Gateway 14.1-73.37 and later, 13.1-64.23 and later, with corresponding

Read More »

Dallas Fed survey: More than one in five firms plan to grow capex in 2027

The share of exploration and production (E&P) companies planning to add to their capital spending in 2027 versus this year has grown to 22% from 10% in June, a new Federal Reserve Bank of Dallas survey shows. Of the more than 80 E&P leaders in Texas, northern Louisiana, and southern New Mexico who responded to the latest Dallas Fed Energy Survey earlier this month, a third said their oil production has increased over the past 3 months and only 1 in 8 said they’re pumping less oil. On the capex side, 46% said their spending this quarter was up from this year’s second quarter. Both of those data points were down slightly from the Fed’s June poll. What appears to be changing more substantially on the ground in the Permian basin, Eagle Ford, and other areas in the Dallas Fed’s footprint are expectations about 2027 spending. Only 5% of E&P leaders now expect they’ll trim capex next year while 73% said they’ll keep spending level. Three months ago, those figures were 10% and 81%, respectively. That means 22% of executives now think their capex will climb in 2027 compared to less than 10% 3 months ago. And it suggests that production in the region will climb from here as producers look to take advantage of consistently high prices for their products—even if they’ve retreated from their recent highs. Jon Costello, an analyst at HFI Research, said an industry response—with Texas firms in the vanguard—to higher prices similar to how it recovered starting in late 2016 would grow total US production more than 4% to about 14.4 million b/d.

Read More »

EIA: US crude oil inventories up 900,000 bbl

US crude oil inventories for the week ended Sept. 25, excluding the Strategic Petroleum Reserve, increased by 900,000 bbl from the previous week, according to data from the US Energy Information Administration (EIA). At 427.3 million bbl, US crude oil inventories are 2% above the 5-year average for this time of year, the EIA report indicated. Gasoline output averaged 9.5 million b/d, and distillate production decreased to 5.0 million b/d. Propane-propylene inventories increased 1.8 million bbl, 20% above the 5-year average. Total commercial petroleum inventories decreased by 7 million bbl for the week. Distillate inventories decreased 2.3 million barrels, 14% below the five-year average. US crude oil refinery inputs averaged 16.3 million b/d for the week ended Sept. 25, which was 554,000 b/d less than the previous week’s average. Refineries operated at 92.5% of capacity. Crude oil imports decreased 179,000 million b/d to 5.7 million b/d. The 4-week average of 6.4 million b/d is 4.8% above the year-ago level. Gasoline imports averaged 500,000 b/d; distillate imports averaged 153,000 b/d. Over the past four weeks, total product supplied averaged 20.8 million b/d, up 2.1% year over year. The 4-week average for gasoline product supplied increased 0.3% year over year to 8.7 million b/d, while the 4-week average for distillate product supplied increased 5.2% to 3.8 million b/d. The 4-week average for jet fuel product supplied increased 6.5% year over year.  

Read More »

INEOS begins commercial CCS at Project Greensand

INEOS Energy and its partners Harbour Energy PLC and Nordsøfonden AS have begun commercial operations at Project Greensand, the European Union’s (EU) first full-scale site for transport and permanent offshore storage of CO2. The CO2 will come mainly from Danish biomethane plants. Once captured, the CO2 is liquefied, sent by truck to a dedicated CO2 terminal at Port Esbjerg, Denmark, shipped aboard the purpose-built CO2 carrier Carbon Destroyer 1, and injected into the Nini West reservoir in the Danish North Sea. Nini West is a depleted oil field 250 km offshore, and roughly 1,800 m below the seabed in about 200-ft water depths. The initial commercial phase provides storage capacity of up to 400,000 tonnes/year (tpy) of CO2, with plans to expand to 4-8 million tpy as demand increases. The EU is working to meet carbon capture and storage targets of 50 million tpy by 2030, rising to 250–280 million tpy by 2040, but remains far from that scale.

Read More »

Morningstar DBRS: Global diesel squeeze boosts US refiners

US refiners are benefiting from a tightening global diesel market as disruptions in the Middle East and Russia constrain supply, lift crack spreads, and keep refinery utilization near capacity, according to Morningstar DBRS. Morningstar DBRS said the current market is creating a strong but likely temporary earnings and cash-flow tailwind for US refiners. High utilization, low inventories, and elevated diesel margins are supporting operating cash flow and EBITDA, although the benefit could fade if geopolitical disruptions ease. Global diesel supply has tightened since the start of the Iran war as refinery outages, lower crude runs, and constraints on product exports through the Strait of Hormuz reduced Middle East supply. Saudi Arabia and Kuwait diesel exports were down about 40% year over year in July. Russia, meanwhile, extended restrictions on most diesel exports into October. Russia exported more than 780,000 b/d of diesel in 2025, just under 10% of global exports. US refiners have increased output and exports to help fill the gap. Distillate production averaged 5.1 million b/d during January-August, the highest since 2019. Refinery utilization is already near capacity in several regions, leaving limited room for further increases in output. PADDs 2 and 4 are operating at or near capacity, supported by discounted Canadian crude, strong diesel export demand, and agricultural and rural consumption. Together, the two regions account for more than 27% of US refining capacity and have an average distillate yield of 32%, DBRS said. On the Gulf Coast, PADD 3 refinery utilization exceeded 98% in September. More than half of US refining capacity is concentrated in PADD 3, where complex refineries serve both export markets and other US regions. The stronger operating environment is translating into higher refining margins. The US Gulf Coast ultra-low-sulfur diesel premium over crude has risen to its highest level this year,

Read More »

US seeks to release another 40 million bbl from SPR despite low inventory levels

The Trump administration Sep. 29 said it would put another 40 million bbl of crude from the Strategic Petroleum Reserve (SPR) into the market even as the emergency stockpile has fallen to its lowest level in more than four decades, raising fresh questions about how much of a buffer remains should another major supply disruption occur. The SPR held 283.8 million bbl as of the week ended Sept. 25. If companies take all 40 million bbl before they return replacement crude, the SPR’s physical inventory would temporarily fall to about 244 million bbl, before accounting for other inventory changes. That would put the stockpile below the 252-million-bbl statutory threshold that applies to certain limited SPR drawdowns. The threshold does not apply to the Department of Energy (DOE)’s exchange authority. SPR risks, exchange demand uncertain The Government Accountability Office (GAO) in May warned that the SPR’s ability to meet future drawdown and fill demands faced risks from aging infrastructure, maintenance backlogs, and low inventory levels. “The SPR’s operational capability is at risk,” the report noted, saying that as of last December, when inventories were over 410 million bbl, the SPR could withdraw oil at only 61% of its design rate and refill the reserve at 56% of its design rate. More than a quarter of the inventory was unavailable for drawdown at the time because of construction and cavern outages. GAO also warned that additional inventory declines from emergency releases could further limit the SPR’s drawdown capability. There is no guarantee that companies will take all 40 million barrels offered under the latest exchange. DOE offered the same amount in June, but only one company agreed to borrow about 500,000 bbl. The limited interest followed concerns among oil traders that the exchange’s repayment premiums and crude-quality requirements could make the SPR

Read More »

Plains names Liollio to succeed Chandler as EVP, COO

Plains All American Pipeline LP and Plains GP Holdings have appointed Dean Liollio to serve as executive vice-president and chief operating officer effective Oct. 2, 2026. Liollio will succeed Chris Chandler, who is resigning from Plains to pursue other interests, the company said in a release Sept. 29. Chandler joined Plains in 2018 after previously serving in leadership roles at Phillips 66. Liollio previously served as senior vice-president, special projects, prior to his appointment as executive vice-president and COO. Prior, the served as president of Plains Midstream Canada from 2020 until 2024, as president of PAA Natural Gas Storage from 2008 until 2020 and as president of Plains Gas Solutions from 2016 until 2020. Prior to joining Plains in 2008, Liollio held a number of executive roles including serving as president, ceief executive officer and director of EnergySouth Inc. and as president and COO of Centerpoint’s natural gas distribution operations across a five-state area.

Read More »

LiquidStack unveils liquid-cooling platform targeting AI data centers

“Operators need cooling infrastructure that can adapt as GPU platforms and rack densities evolve,” said Scott Smith, general manager of LiquidStack, in a statement. “CDU 2.X combines the performance and flexibility customers need today with the headroom to prepare for what comes next, allowing them to configure cooling around their facility and deployment strategy rather than designing the facility around the CDU.” A key feature is the platform’s support for different deployment configurations, including end-of-row and rack-adjacent installations. The idea is to give data center operators more flexibility in how they place their equipment with increasingly high thermal loads. The system offers configurable control-valve, power-feed and redundancy options, including dual-feed A/B configurations and automatic transfer switch support. These features allow operators to tailor the CDU to different facility architectures and resiliency requirements.

Read More »

If Apple returns to enterprise server game (with help from Nvidia), what market could it target?

Nvidia introduced NVLink Fusion as part of a broader effort to make its interconnect technology available to companies developing custom AI processors. The approach allows third-party silicon to be incorporated into Nvidia-oriented data center architectures, potentially extending Nvidia’s influence beyond its own GPUs. A similar strategy is already being pursued with inference-chip developer d-Matrix. Apple also already operates specialized servers for its Private Cloud Compute system, which handles Apple Intelligence workloads that require processing beyond the user’s device. Scaling that infrastructure, however, reportedly creates bandwidth, cost, and performance challenges. A commercial AI server would allow Apple to extend its silicon strategy into the data center while giving customers direct access to Apple processors rather than Apple’s own cloud services.

Read More »

From Coal to Compute: How Pennsylvania Is Rebuilding Power for AI Data Centers

Western Pennsylvania is becoming a test bed for one of the most consequential changes underway in data center development: the migration from simply finding grid capacity to building the power supply along with the data center. Two projects illustrate that transition particularly well. Aligned Data Centers is advancing the roughly $10 billion, 2-GW Project Phoenix campus at the former Bruce Mansfield coal-fired power station site in Shippingport, Beaver County. About 60 miles to the east, the former Homer City Generating Station is being transformed into the Homer City Energy Campus, centered on as much as 4.4 GW of new natural-gas generation and a proposed Amazon Web Services campus that could ultimately include 39 data center buildings. The projects share a number of similarities. Both reuse former coal-generation properties. Both already possess much of the infrastructure that greenfield data center developers spend years trying to obtain: high-voltage transmission, industrial zoning, water infrastructure, pipeline access, large parcels and proximity to the Marcellus and Utica natural-gas fields.But technically and commercially, they are also very different. Project Phoenix is essentially a data-center-led development using behind-the-meter generation, supplemented by a separate plan to repower the former coal station. While Homer City is the reverse: a power-generation project being constructed first, with the hyperscale data center development forming around the available power supply. This isn’t a competition, but it is a comparison on speed of delivery and the issues both development models face. Project Phoenix: Aligned Moves Into Pennsylvania Aligned calls Shippingport its first Pennsylvania campus and a regional flagship. The company formally broke ground on Project Phoenix on September 10, 2026, describing it as a 2-GW campus spanning three data center facilities and representing roughly $10 billion in regional investment. Aligned estimates the development will support approximately 3,000 construction jobs and 640 full-time jobs in

Read More »

OpenStack Hibiscus adds DNS security features and confidential computing to the open-source cloud platform

Type-5 support is aimed at data center integration. “It lets the tenant network prefixes be advertised directly into the physical fabric, so that really that’s basically how modern data center networks are built,” Carrez explained. “So really helps OpenStack fit into those environments without extra gateway layers that we’ve seen in use before.” Routable tenant addresses: The OVN BGP integration gains a route leaking option. An operator turns it on with the leak_routes attribute of a subnet. The extended OVN features follow the same approach. “OVN BGP features that let the tenant addresses be routable directly from the underlay, and that again is exposing how modern data centers are built directly into OpenStack,” Carrez said. Lower memory use: In deployments that use Open vSwitch, a monitoring daemon tracked keepalived state changes for HA routers. Hibiscus replaces the daemon with a shell script. The change applies to every HA router, so the savings add up across a deployment. “The 15 times reduction in memory footprint for the high availability router monitoring is, I think, really interesting,” Carrez said.

Read More »

DCF Trends Summit: ON.energy’s Asser Elsamahy – Using AI UPS Systems to Tame AI Load Swings

The power challenge surrounding artificial intelligence is increasingly about more than finding enough megawatts. AI data centers can also introduce rapid changes in electricity demand as large clusters of accelerators ramp workloads up and down. Those swings create a different kind of infrastructure problem: how to serve highly dynamic compute loads without passing that volatility directly onto the electric grid. That challenge is helping move battery energy storage deeper into data center power architecture. In an onsite podcast interview recorded live at the Data Center Frontier Trends Summit 2026, DCF Contributing Editor Doug Black spoke with Asser Elsamahy, P.E., vice president of engineering at ON.energy, about the emerging role of battery-based power quality infrastructure for AI data centers. Elsamahy said battery power systems themselves are hardly new. Energy storage has been deployed at gigawatt scale around the world for roughly two decades. What is new is the way the technology is being adapted to the operating characteristics of large AI facilities. “They’re new to the data center industry, but they’re not necessarily new in the market,” Elsamahy said. “They’ve been deployed at gigawatt scale already, multiple gigawatts all over the world.” The difference now is the load. Major swings in AI computing demand can create additional stress for grid operators already confronting rapid growth in large-load interconnection requests. Elsamahy said that dynamic is accelerating interest in energy storage as a way to manage the interface between AI infrastructure and the grid. From Battery Storage to an “AI UPS” ON.energy’s approach is built around what the company calls an AI UPS, or medium-voltage uninterruptible power supply. The architecture differs from the parallel battery energy storage system, or BESS, configuration commonly used for standalone grid storage. ON instead uses a double-conversion design with two sets of inverters. One inverter set faces the

Read More »

Q3 Executive Roundtable Recap

For Data Center Frontier’s Q3 2026 Executive Roundtable, three industry leaders examined a question increasingly central to the AI infrastructure buildout: What happens when data centers are asked to become larger, denser and faster at the same time? Across three discussions, a consistent theme emerged. AI is not simply increasing the amount of infrastructure required to support the modern data center. It is exposing assumptions that were easier to tolerate at lower densities, expanding the boundaries of what operators must consider mission-critical, and making the interaction between systems increasingly important to overall resilience. That begins with density. As the value and power concentrated in individual racks rises, traditional approaches to redundancy, monitoring and risk mitigation can leave less room for error. Resilience can no longer be measured simply by installed capacity or the presence of backup equipment. Operators increasingly need to understand how electrical, thermal and control systems behave together under dynamic AI workloads — and how quickly the facility can respond and recover when something goes wrong. The same shift is broadening the definition of critical infrastructure. Power generation, UPS systems and network connectivity remain fundamental, but energy storage, liquid cooling, leak detection, controls, monitoring and the interfaces connecting them are becoming part of the same reliability equation. A component can perform exactly as designed while the larger system still fails if coordination, communications or control logic break down. And all of this is happening while the market is demanding faster deployment. Standardization, modular construction, factory integration and earlier modeling can legitimately compress project schedules. But the Q3 panelists drew a clear distinction between eliminating unnecessary time and eliminating rigor. As infrastructure becomes more tightly coupled, commissioning, integrated systems testing, operational visibility and system-level validation may need to become more thorough precisely because projects are moving faster. Taken together,

Read More »

Microsoft will invest $80B in AI data centers in fiscal 2025

And Microsoft isn’t the only one that is ramping up its investments into AI-enabled data centers. Rival cloud service providers are all investing in either upgrading or opening new data centers to capture a larger chunk of business from developers and users of large language models (LLMs).  In a report published in October 2024, Bloomberg Intelligence estimated that demand for generative AI would push Microsoft, AWS, Google, Oracle, Meta, and Apple would between them devote $200 billion to capex in 2025, up from $110 billion in 2023. Microsoft is one of the biggest spenders, followed closely by Google and AWS, Bloomberg Intelligence said. Its estimate of Microsoft’s capital spending on AI, at $62.4 billion for calendar 2025, is lower than Smith’s claim that the company will invest $80 billion in the fiscal year to June 30, 2025. Both figures, though, are way higher than Microsoft’s 2020 capital expenditure of “just” $17.6 billion. The majority of the increased spending is tied to cloud services and the expansion of AI infrastructure needed to provide compute capacity for OpenAI workloads. Separately, last October Amazon CEO Andy Jassy said his company planned total capex spend of $75 billion in 2024 and even more in 2025, with much of it going to AWS, its cloud computing division.

Read More »

John Deere unveils more autonomous farm machines to address skill labor shortage

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Self-driving tractors might be the path to self-driving cars. John Deere has revealed a new line of autonomous machines and tech across agriculture, construction and commercial landscaping. The Moline, Illinois-based John Deere has been in business for 187 years, yet it’s been a regular as a non-tech company showing off technology at the big tech trade show in Las Vegas and is back at CES 2025 with more autonomous tractors and other vehicles. This is not something we usually cover, but John Deere has a lot of data that is interesting in the big picture of tech. The message from the company is that there aren’t enough skilled farm laborers to do the work that its customers need. It’s been a challenge for most of the last two decades, said Jahmy Hindman, CTO at John Deere, in a briefing. Much of the tech will come this fall and after that. He noted that the average farmer in the U.S. is over 58 and works 12 to 18 hours a day to grow food for us. And he said the American Farm Bureau Federation estimates there are roughly 2.4 million farm jobs that need to be filled annually; and the agricultural work force continues to shrink. (This is my hint to the anti-immigration crowd). John Deere’s autonomous 9RX Tractor. Farmers can oversee it using an app. While each of these industries experiences their own set of challenges, a commonality across all is skilled labor availability. In construction, about 80% percent of contractors struggle to find skilled labor. And in commercial landscaping, 86% of landscaping business owners can’t find labor to fill open positions, he said. “They have to figure out how to do

Read More »

2025 playbook for enterprise AI success, from agents to evals

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More 2025 is poised to be a pivotal year for enterprise AI. The past year has seen rapid innovation, and this year will see the same. This has made it more critical than ever to revisit your AI strategy to stay competitive and create value for your customers. From scaling AI agents to optimizing costs, here are the five critical areas enterprises should prioritize for their AI strategy this year. 1. Agents: the next generation of automation AI agents are no longer theoretical. In 2025, they’re indispensable tools for enterprises looking to streamline operations and enhance customer interactions. Unlike traditional software, agents powered by large language models (LLMs) can make nuanced decisions, navigate complex multi-step tasks, and integrate seamlessly with tools and APIs. At the start of 2024, agents were not ready for prime time, making frustrating mistakes like hallucinating URLs. They started getting better as frontier large language models themselves improved. “Let me put it this way,” said Sam Witteveen, cofounder of Red Dragon, a company that develops agents for companies, and that recently reviewed the 48 agents it built last year. “Interestingly, the ones that we built at the start of the year, a lot of those worked way better at the end of the year just because the models got better.” Witteveen shared this in the video podcast we filmed to discuss these five big trends in detail. Models are getting better and hallucinating less, and they’re also being trained to do agentic tasks. Another feature that the model providers are researching is a way to use the LLM as a judge, and as models get cheaper (something we’ll cover below), companies can use three or more models to

Read More »

OpenAI’s red teaming innovations define new essentials for security leaders in the AI era

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More OpenAI has taken a more aggressive approach to red teaming than its AI competitors, demonstrating its security teams’ advanced capabilities in two areas: multi-step reinforcement and external red teaming. OpenAI recently released two papers that set a new competitive standard for improving the quality, reliability and safety of AI models in these two techniques and more. The first paper, “OpenAI’s Approach to External Red Teaming for AI Models and Systems,” reports that specialized teams outside the company have proven effective in uncovering vulnerabilities that might otherwise have made it into a released model because in-house testing techniques may have missed them. In the second paper, “Diverse and Effective Red Teaming with Auto-Generated Rewards and Multi-Step Reinforcement Learning,” OpenAI introduces an automated framework that relies on iterative reinforcement learning to generate a broad spectrum of novel, wide-ranging attacks. Going all-in on red teaming pays practical, competitive dividends It’s encouraging to see competitive intensity in red teaming growing among AI companies. When Anthropic released its AI red team guidelines in June of last year, it joined AI providers including Google, Microsoft, Nvidia, OpenAI, and even the U.S.’s National Institute of Standards and Technology (NIST), which all had released red teaming frameworks. Investing heavily in red teaming yields tangible benefits for security leaders in any organization. OpenAI’s paper on external red teaming provides a detailed analysis of how the company strives to create specialized external teams that include cybersecurity and subject matter experts. The goal is to see if knowledgeable external teams can defeat models’ security perimeters and find gaps in their security, biases and controls that prompt-based testing couldn’t find. What makes OpenAI’s recent papers noteworthy is how well they define using human-in-the-middle

Read More »

Introducing SynthID Bio

Strengthening biosecurity and information integrityBiosecurity relies on layered defenses – think of it like a “Swiss cheese” defense model, where multiple independent safety measures work

Read More »