LettuceDetect: A Hallucination Detection Framework for RAG Applications

Stay Ahead, Stay ONMINE

LettuceDetect: A Hallucination Detection Framework for RAG Applications

Originally published on HuggingFace TL;DR We present LettuceDetect, a lightweight hallucination detector for Retrieval-Augmented Generation (RAG) pipelines. It is an encoder-based model built on ModernBERT, released under the MIT license with ready-to-use Python packages and pretrained models. What: LettuceDetect is a token-level detector that flags unsupported segments in LLM answers. 🥬 How: Trained on RAGTruth (18k examples), leveraging ModernBERT for context lengths up to 4k tokens. 🚀 Why: It addresses (1) the context-window limits in prior encoder-only models, and (2) the high compute costs of LLM-based detectors. ⚖️ Highlights: Beats prior encoder-based models (e.g., Luna) on RAGTruth. ✅ Surpasses fine-tuned Llama-2-13B [2] at a fraction of the size, and is highly efficient at inference. ⚡️ Entirely open-source with an MIT license. 🔓 LettuceDetect keeps your RAG framework fresh by spotting rotten parts of your LLM’s outputs. 😊 Quick links Why LettuceDetect? Large Language Models (LLMs) have made considerable advancements in NLP tasks, like GPT-4 [4], the Llama-3 models [5], or Mistral [6] (and many more). Despite the success of LLMs, Hallucinations remain a key obstacle deploying LLMs in high-stakes scenarios (such as in healthcare or legal) [7,8]. Retrieval-Augmented Generation (RAG) attempts to mitigate hallucinations by grounding an LLM’s responses in retrieved documents, providing external knowledge that the model can reference [9]. But even though RAG is a powerful method to reduce hallucinations, LLMs still suffer from hallucinations in these settings [1]. Hallucinations are information in the output that is nonsensical, factually incorrect, or inconsistent with the retrieved context [8]. Ji et al. [10] categorizes hallucinations into: Intrinsic hallucinations: Stemming from the model’s preexisting internal knowledge. Extrinsic hallucinations: Occurring when the answer conflicts with the context or references provided While RAG approaches can mitigate intrinsic hallucinations, they are not immune to extrinsic hallucinations. Sun et al. [11] showed that models tend to prioritize their intrinsic knowledge over the external context. As LLMs remain prone to hallucinations, their applications in critical domains e.g. medical or legal, can be still flawed. Current solutions for hallucination detection Current solutions for hallucination detection can be categorized into different categories based on the approach they take: Prompt-based detectors These methods (e.g., RAGAS, Trulens, ARES) typically leverage zero-shot or few-shot prompts to detect hallucinations. They often rely on large LLMs (like GPT-4) and employ strategies such as SelfCheckGPT [12], LM vs. LM [13], or Chainpoll [14]. While often effective, they can be computationally expensive due to repeated LLM calls. Fine-tuned LLM detectors Large models (e.g., Llama-2, Llama-3) can be fine-tuned for hallucination detection [1,15]. This can yield high accuracy (as shown by the RAGTruth authors using Llama-2-13B or the RAG-HAT work on Llama-3-8B) but is resource-intensive to train and deploy. Inference costs also tend to be high due to their size and slower speeds. Encoder-based detectors Models like Luna [2] rely on a BERT-style encoder (often limited to 512 tokens) for token-level classification. These methods are generally more efficient than running a full LLM at inference but are constrained by short context windows and attention mechanisms optimized for smaller inputs. ModernBERT for long context ModernBERT [3] is a drop-in replacement for BERT and is a state-of-the-art encoder-only transformers architecture that incorporates several modern design improvements over the original BERT model such as it uses Rotary Positional Embeddings (RoPe) to handle sequences of up to 8,192 tokens, unpadding optimization to eliminate wasted computation on padding tokens, and GeGLU activation layers for enhanced expressiveness and alternating attention for more efficient attention computation. LettuceDetect capitalizes on ModernBERT’s extended context window to build a token-level classifier for hallucination detection. This approach sidesteps many limitations of older BERT-based models (e.g., short context bounds) and avoids the inference overhead of large LLM-based detectors. Our experiments show that LettuceDetect outperforms other encoder-based systems while remaining competitive with fine-tuned LLM detectors at a fraction of their computational cost. Data RAGTruth is the first large-scale open-source (MIT License) benchmark specifically designed to evaluate hallucination detection in Retrieval-Augmented Generation (RAG) settings. It contains 18,000 annotated examples spanning multiple tasks: Question answering (QA): Sampled from the MS MARCO dataset, where up to three documents are retrieved for each question, and an LLM is prompted to produce an answer. Data-to-text generation: Based on the Yelp Open Dataset, where LLMs generate reviews for sampled businesses. News summarization: Uses random documents from CNN/DailyMail, prompting LLMs to provide summaries. Each example in RAGTruth includes: A question and prompt, One or more reference passages or context documents, An LLM-generated response, Human annotations at the span level indicating whether any part of the answer is factually unsupported or contradictory to the reference passages. Below is an illustrative sample from RAGTruth, showing an instance of hallucination: QUESTION How to explain quartiles? REFERENCE Passage 2:Second quartile (Q2) which is more commonly known as median splits the data in half (50%).Median divides the data into a lower half and an upper half.Third quartile (Q3), also known as upper quartile, splits lowest 75% (or highest 25%) of data.(…) RESPONSE The first quartile (Q1) splits the lowest 25% of the data, while the second quartile (Q2) splits the data into two equal halves, with the median being the middle value of the lower half. Finally, the third quartile (Q3) splits the highest 75% of the data.(…) ANNOTATION Span: “highest 75%” Type: Evident Conflict Reason: Original text: “lowest 75% (or highest 25%)” Generative text: “highest 75%” This example showcases a RAGTruth annotation involving specific spans in the answer that conflict with the reference passages, providing both the type of hallucination (e.g., Evident Conflict) and a human-readable rationale. Our method only uses binary annotations and does not use the type of hallucination annotations. Method Image by Author A high-level depiction of LettuceDetect. Here, an example Question, Context, and Answer triplet is processed. First, the text is tokenized, after which LettuceDetect performs token-level classification. Tokens from both the question and context are masked (indicated by the red line in the figure) to exclude them from the loss function. Each token in the answer receives a probability indicating whether it is hallucinated or supported. For span-level detection, we merge consecutive tokens with hallucination probabilities above 0.5 into a single predicted span. We train ModernBERT-base and ModernBERT-large variants as token-classification models on the RAGTruth dataset. The input to the model is a concatenation of Context, Question, and Answer segments, with specialized tokens ([CLS]) (for the context) and ([SEP]) (as separators). We limit the sequence length to 4,096 tokens for computational feasibility, though ModernBERT can theoretically handle up to 8,192 tokens. Tokenization and data processing Tokenizer: We employ AutoTokenizer from the Transformers library to handle subword Tokenization, inserting [CLS] and [SEP] appropriately. Labeling: Context/question tokens are masked (i.e., assigned a label of -100 in PyTorch) so that they do not contribute to the loss. Each answer token receives a label of 0 (supported) or 1 (hallucinated). Model architecture Our models build on Hugging Face’s AutoModelForTokenClassification, using ModernBERT as the encoder and a classification head on top. Unlike some previous encoder-based approaches (e.g., ones pre-trained on NLI tasks), our method uses only ModernBERT with no additional pretraining stage. Training configuration Optimizer: AdamW, with a learning rate of 1 * 10^-5 and weight decay of 0.01. Hardware: Single NVIDIA A100 GPU. Epochs: 6 total training epochs. Batching: Batch size of 8, Data loading with PyTorch DataLoader (shuffling enabled), Dynamic padding via DataCollatorForTokenClassification to handle variable-length sequences efficiently. During training, we monitor token-level F1 scores on a validation split, saving checkpoints using the safetensors format. Once training is complete, we upload the best-performing models to Hugging Face for public access. At inference time, the model outputs a probability of hallucination for each token in the answer. We aggregate consecutive tokens exceeding a 0.5 threshold to produce span-level predictions, indicating exactly which segments of the answer are likely to be hallucinated. The figure above illustrates this workflow. Next, we provide a more detailed evaluation of the model’s performance. Results We evaluate our models on the RAGTruth test set across all task types (Question Answering, Data-to-Text, and Summarization). For each example, RAGTruth includes manually annotated spans indicating hallucinated content. Example-level results We first assess the example-level question: Does the generated answer contain any hallucination at all? Our large model (lettucedetect-large-v1) attains an overall F1 score of 79.22%, surpassing: GPT-4 (63.4%), Luna (65.4%) (the previous state of the art encoder-based model), Fine-tuned Llama-2-13B (78.7%) as presented in the RAGTruth paper. It is second only to the fine-tuned Llama-3-8B from the RAG-HAT paper [15] (83.9%), but LettuceDetect is significantly smaller and faster to run. Meanwhile, our base model (lettucedetect-base-v1) remains highly competitive while using fewer parameters. Image by Author Above is a comparison table illustrating how LettuceDetect aligns against both prompt-based methods (e.g., GPT-4) and alternative encoder-based solutions (e.g., Luna). Overall, lettucedetect-large-v1 and lettucedect-base-v1 are very performant models, while being very effective in inference settings. Span-level results Beyond detecting if an answer contains hallucinations, we also examine LettuceDetect’s ability to identify the exact spans of unsupported content. Here, LettuceDetect achieves state-of-the-art results among models that have reported span-level performance, substantially outperforming the fine-tuned Llama-2-13B model from the RAGTruth paper [1] and other baselines. Image by Author Most methods, like RAG-HAT [15], do not report span-level metrics, so we do not compare to them here. Inference efficiency Both lettucedetect-base-v1 and lettucedetect-large-v1 require fewer parameters than typical LLM-based detectors (e.g., GPT-4 or Llama-3-8B) and can process 30–60 examples per second on a single NVIDIA A100 GPU. This makes them practical for industrial workloads, real-time user-facing systems, and resource-constrained environments. Overall, these results show that LettuceDetect has a good balance: it achieves near state-of-the-art accuracy at a fraction of the size and cost compared to large LLM-based judges, while offering precise, token-level hallucination detection. Get going Install the package: pip install lettucedetect Then, you can use the package as follows: from lettucedetect.models.inference import HallucinationDetector # For a transformer-based approach: detector = HallucinationDetector( method=”transformer”, model_path=”KRLabsOrg/lettucedect-base-modernbert-en-v1″ ) contexts = [“France is a country in Europe. The capital of France is Paris. The population of France is 67 million.”,] question = “What is the capital of France? What is the population of France?” answer = “The capital of France is Paris. The population of France is 69 million.” # Get span-level predictions indicating which parts of the answer are considered hallucinated. predictions = detector.predict(context=contexts, question=question, answer=answer, output_format=”spans”) print(“Predictions:”, predictions) # Predictions: [{‘start’: 31, ‘end’: 71, ‘confidence’: 0.9944414496421814, ‘text’: ‘ The population of France is 69 million.’}] Conclusion We introduced LettuceDetect, a lightweight and efficient framework for hallucination detection in RAG systems. By utilizing ModernBERT’s extended context capabilities, our models achieve strong performance on the RAGTruth benchmark while retaining high inference efficiency. This work lays the groundwork for future research directions, such as expanding to additional datasets, supporting multiple languages, and exploring more advanced architectures. Even at this stage, LettuceDetect demonstrates that effective hallucination detection can be achieved using lean, purpose-built encoder-based models. Citation If you find this work useful, please cite it as follows: @misc{Kovacs:2025, title={LettuceDetect: A Hallucination Detection Framework for RAG Applications}, author={Ádám Kovács and Gábor Recski}, year={2025}, eprint={2502.17125}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2502.17125}, } Also, if you use our code, please don’t forget to give us a star ⭐ on our GitHub repository here. References [1] Niu et al., 2024, RAGTruth: A Dataset for Hallucination Detection in Retrieval-Augmented Generation [2] Luna: A Simple and Effective Encoder-Based Model for Hallucination Detection in Retrieval-Augmented Generation [3] ModernBERT: A Modern BERT Model for Long-Context Processing [4] GPT-4 report [5] Llama-3 report [6] Mistral 7B [7] Kaddour et al., 2023, Challenges and Applications of Large Language Models [8] Huang et al., 2025, A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions [9] Gao et al., 2024, Retrieval-Augmented Generation for Large Language Models: A Survey [10] Ji et al., 2023, Survey of Hallucination in Natural Language Generation [11] Sun et al., 2025, ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability [12] Manakul et al., 2023, SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models [13] Cohen et al., 2023, LM vs LM: Detecting Factual Errors via Cross Examination [14] Friel et al., 2023, Chainpoll: A high efficacy method for LLM hallucination detection [15] Song et al., 2024, RAG-HAT: A Hallucination-Aware Tuning Pipeline for {LLM} in Retrieval-Augmented Generation [16] Devlin et al., 2019, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Originally published on HuggingFace

TL;DR

We present LettuceDetect, a lightweight hallucination detector for Retrieval-Augmented Generation (RAG) pipelines. It is an encoder-based model built on ModernBERT, released under the MIT license with ready-to-use Python packages and pretrained models.

What: LettuceDetect is a token-level detector that flags unsupported segments in LLM answers. 🥬
How: Trained on RAGTruth (18k examples), leveraging ModernBERT for context lengths up to 4k tokens. 🚀
Why: It addresses (1) the context-window limits in prior encoder-only models, and (2) the high compute costs of LLM-based detectors. ⚖️

Highlights:
- Beats prior encoder-based models (e.g., Luna) on RAGTruth. ✅
- Surpasses fine-tuned Llama-2-13B [2] at a fraction of the size, and is highly efficient at inference. ⚡️
- Entirely open-source with an MIT license. 🔓

LettuceDetect keeps your RAG framework fresh by spotting rotten parts of your LLM’s outputs. 😊

Quick links

Why LettuceDetect?

Large Language Models (LLMs) have made considerable advancements in NLP tasks, like GPT-4 [4], the Llama-3 models [5], or Mistral [6] (and many more). Despite the success of LLMs, Hallucinations remain a key obstacle deploying LLMs in high-stakes scenarios (such as in healthcare or legal) [7,8].

Retrieval-Augmented Generation (RAG) attempts to mitigate hallucinations by grounding an LLM’s responses in retrieved documents, providing external knowledge that the model can reference [9]. But even though RAG is a powerful method to reduce hallucinations, LLMs still suffer from hallucinations in these settings [1]. Hallucinations are information in the output that is nonsensical, factually incorrect, or inconsistent with the retrieved context [8]. Ji et al. [10] categorizes hallucinations into:

Intrinsic hallucinations: Stemming from the model’s preexisting internal knowledge.
Extrinsic hallucinations: Occurring when the answer conflicts with the context or references provided

While RAG approaches can mitigate intrinsic hallucinations, they are not immune to extrinsic hallucinations. Sun et al. [11] showed that models tend to prioritize their intrinsic knowledge over the external context. As LLMs remain prone to hallucinations, their applications in critical domains e.g. medical or legal, can be still flawed.

Current solutions for hallucination detection

Current solutions for hallucination detection can be categorized into different categories based on the approach they take:

Prompt-based detectors These methods (e.g., RAGAS, Trulens, ARES) typically leverage zero-shot or few-shot prompts to detect hallucinations. They often rely on large LLMs (like GPT-4) and employ strategies such as SelfCheckGPT [12], LM vs. LM [13], or Chainpoll [14]. While often effective, they can be computationally expensive due to repeated LLM calls.
Fine-tuned LLM detectors Large models (e.g., Llama-2, Llama-3) can be fine-tuned for hallucination detection [1,15]. This can yield high accuracy (as shown by the RAGTruth authors using Llama-2-13B or the RAG-HAT work on Llama-3-8B) but is resource-intensive to train and deploy. Inference costs also tend to be high due to their size and slower speeds.
Encoder-based detectors Models like Luna [2] rely on a BERT-style encoder (often limited to 512 tokens) for token-level classification. These methods are generally more efficient than running a full LLM at inference but are constrained by short context windows and attention mechanisms optimized for smaller inputs.

ModernBERT for long context

ModernBERT [3] is a drop-in replacement for BERT and is a state-of-the-art encoder-only transformers architecture that incorporates several modern design improvements over the original BERT model such as it uses Rotary Positional Embeddings (RoPe) to handle sequences of up to 8,192 tokens, unpadding optimization to eliminate wasted computation on padding tokens, and GeGLU activation layers for enhanced expressiveness and alternating attention for more efficient attention computation.

LettuceDetect capitalizes on ModernBERT’s extended context window to build a token-level classifier for hallucination detection. This approach sidesteps many limitations of older BERT-based models (e.g., short context bounds) and avoids the inference overhead of large LLM-based detectors. Our experiments show that LettuceDetect outperforms other encoder-based systems while remaining competitive with fine-tuned LLM detectors at a fraction of their computational cost.

Data

RAGTruth is the first large-scale open-source (MIT License) benchmark specifically designed to evaluate hallucination detection in Retrieval-Augmented Generation (RAG) settings. It contains 18,000 annotated examples spanning multiple tasks:

Question answering (QA): Sampled from the MS MARCO dataset, where up to three documents are retrieved for each question, and an LLM is prompted to produce an answer.
Data-to-text generation: Based on the Yelp Open Dataset, where LLMs generate reviews for sampled businesses.
News summarization: Uses random documents from CNN/DailyMail, prompting LLMs to provide summaries.

Each example in RAGTruth includes:

A question and prompt,
One or more reference passages or context documents,
An LLM-generated response,
Human annotations at the span level indicating whether any part of the answer is factually unsupported or contradictory to the reference passages.

Below is an illustrative sample from RAGTruth, showing an instance of hallucination:

QUESTION

How to explain quartiles?

REFERENCE

Passage 2:
Second quartile (Q2) which is more commonly known as median splits the data in half (50%).
Median divides the data into a lower half and an upper half.
Third quartile (Q3), also known as upper quartile, splits lowest 75% (or highest 25%) of data.
(…)

RESPONSE

The first quartile (Q1) splits the lowest 25% of the data, while the second quartile (Q2) splits the data into two equal halves, with the median being the middle value of the lower half. Finally, the third quartile (Q3) splits the highest 75% of the data.
(…)

ANNOTATION

Span: “highest 75%”
Type: Evident Conflict
Reason:
- Original text: “lowest 75% (or highest 25%)”
- Generative text: “highest 75%”

This example showcases a RAGTruth annotation involving specific spans in the answer that conflict with the reference passages, providing both the type of hallucination (e.g., Evident Conflict) and a human-readable rationale. Our method only uses binary annotations and does not use the type of hallucination annotations.

Method

Diagram of LettuceDetect — Image by Author

A high-level depiction of LettuceDetect. Here, an example Question, Context, and Answer triplet is processed. First, the text is tokenized, after which LettuceDetect performs token-level classification. Tokens from both the question and context are masked (indicated by the red line in the figure) to exclude them from the loss function. Each token in the answer receives a probability indicating whether it is hallucinated or supported. For span-level detection, we merge consecutive tokens with hallucination probabilities above 0.5 into a single predicted span.

We train ModernBERT-base and ModernBERT-large variants as token-classification models on the RAGTruth dataset. The input to the model is a concatenation of Context, Question, and Answer segments, with specialized tokens ([CLS]) (for the context) and ([SEP]) (as separators). We limit the sequence length to 4,096 tokens for computational feasibility, though ModernBERT can theoretically handle up to 8,192 tokens.

Tokenization and data processing

Tokenizer: We employ AutoTokenizer from the Transformers library to handle subword Tokenization, inserting [CLS] and [SEP] appropriately.
Labeling:
- Context/question tokens are masked (i.e., assigned a label of -100 in PyTorch) so that they do not contribute to the loss.
- Each answer token receives a label of 0 (supported) or 1 (hallucinated).

Model architecture

Our models build on Hugging Face’s AutoModelForTokenClassification, using ModernBERT as the encoder and a classification head on top. Unlike some previous encoder-based approaches (e.g., ones pre-trained on NLI tasks), our method uses only ModernBERT with no additional pretraining stage.

Training configuration

Optimizer: AdamW, with a learning rate of 1 * 10^-5 and weight decay of 0.01.
Hardware: Single NVIDIA A100 GPU.
Epochs: 6 total training epochs.
Batching:
- Batch size of 8,
- Data loading with PyTorch DataLoader (shuffling enabled),
- Dynamic padding via DataCollatorForTokenClassification to handle variable-length sequences efficiently.

During training, we monitor token-level F1 scores on a validation split, saving checkpoints using the safetensors format. Once training is complete, we upload the best-performing models to Hugging Face for public access.

At inference time, the model outputs a probability of hallucination for each token in the answer. We aggregate consecutive tokens exceeding a 0.5 threshold to produce span-level predictions, indicating exactly which segments of the answer are likely to be hallucinated. The figure above illustrates this workflow.

Next, we provide a more detailed evaluation of the model’s performance.

Results

We evaluate our models on the RAGTruth test set across all task types (Question Answering, Data-to-Text, and Summarization). For each example, RAGTruth includes manually annotated spans indicating hallucinated content.

Example-level results

We first assess the example-level question: Does the generated answer contain any hallucination at all? Our large model (lettucedetect-large-v1) attains an overall F1 score of 79.22%, surpassing:

GPT-4 (63.4%),
Luna (65.4%) (the previous state of the art encoder-based model),
Fine-tuned Llama-2-13B (78.7%) as presented in the RAGTruth paper.

It is second only to the fine-tuned Llama-3-8B from the RAG-HAT paper [15] (83.9%), but LettuceDetect is significantly smaller and faster to run. Meanwhile, our base model (lettucedetect-base-v1) remains highly competitive while using fewer parameters.

Comparison table illustrating how LettuceDetect aligns against both prompt-based methods (e.g., GPT-4) and alternative encoder-based solutions (e.g., Luna) — Image by Author

Above is a comparison table illustrating how LettuceDetect aligns against both prompt-based methods (e.g., GPT-4) and alternative encoder-based solutions (e.g., Luna). Overall, lettucedetect-large-v1 and lettucedect-base-v1 are very performant models, while being very effective in inference settings.

Span-level results

Beyond detecting if an answer contains hallucinations, we also examine LettuceDetect’s ability to identify the exact spans of unsupported content. Here, LettuceDetect achieves state-of-the-art results among models that have reported span-level performance, substantially outperforming the fine-tuned Llama-2-13B model from the RAGTruth paper [1] and other baselines.

Most methods, like RAG-HAT [15], do not report span-level metrics, so we do not compare to them here.

Inference efficiency

Both lettucedetect-base-v1 and lettucedetect-large-v1 require fewer parameters than typical LLM-based detectors (e.g., GPT-4 or Llama-3-8B) and can process 30–60 examples per second on a single NVIDIA A100 GPU. This makes them practical for industrial workloads, real-time user-facing systems, and resource-constrained environments.

Overall, these results show that LettuceDetect has a good balance: it achieves near state-of-the-art accuracy at a fraction of the size and cost compared to large LLM-based judges, while offering precise, token-level hallucination detection.

Get going

Install the package:

pip install lettucedetect

Then, you can use the package as follows:

from lettucedetect.models.inference import HallucinationDetector

# For a transformer-based approach:

detector = HallucinationDetector(

    method="transformer", model_path="KRLabsOrg/lettucedect-base-modernbert-en-v1"

)

contexts = ["France is a country in Europe. The capital of France is Paris. The population of France is 67 million.",]

question = "What is the capital of France? What is the population of France?"

answer = "The capital of France is Paris. The population of France is 69 million."

# Get span-level predictions indicating which parts of the answer are considered hallucinated.

predictions = detector.predict(context=contexts, question=question, answer=answer, output_format="spans")

print("Predictions:", predictions)

# Predictions: [{'start': 31, 'end': 71, 'confidence': 0.9944414496421814, 'text': ' The population of France is 69 million.'}]

Conclusion

We introduced LettuceDetect, a lightweight and efficient framework for hallucination detection in RAG systems. By utilizing ModernBERT’s extended context capabilities, our models achieve strong performance on the RAGTruth benchmark while retaining high inference efficiency. This work lays the groundwork for future research directions, such as expanding to additional datasets, supporting multiple languages, and exploring more advanced architectures. Even at this stage, LettuceDetect demonstrates that effective hallucination detection can be achieved using lean, purpose-built encoder-based models.

Citation

If you find this work useful, please cite it as follows:

@misc{Kovacs:2025,
      title={LettuceDetect: A Hallucination Detection Framework for RAG Applications}, 
      author={Ádám Kovács and Gábor Recski},
      year={2025},
      eprint={2502.17125},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2502.17125}, 
}

Also, if you use our code, please don’t forget to give us a star ⭐ on our GitHub repository here.

References

[1] Niu et al., 2024, RAGTruth: A Dataset for Hallucination Detection in Retrieval-Augmented Generation

[2] Luna: A Simple and Effective Encoder-Based Model for Hallucination Detection in Retrieval-Augmented Generation

[3] ModernBERT: A Modern BERT Model for Long-Context Processing

[4] GPT-4 report

[5] Llama-3 report

[6] Mistral 7B

[7] Kaddour et al., 2023, Challenges and Applications of Large Language Models

[8] Huang et al., 2025, A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

[9] Gao et al., 2024, Retrieval-Augmented Generation for Large Language Models: A Survey

[10] Ji et al., 2023, Survey of Hallucination in Natural Language Generation

[11] Sun et al., 2025, ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability

[12] Manakul et al., 2023, SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models

[13] Cohen et al., 2023, LM vs LM: Detecting Factual Errors via Cross Examination

[14] Friel et al., 2023, Chainpoll: A high efficacy method for LLM hallucination detection

[15] Song et al., 2024, RAG-HAT: A Hallucination-Aware Tuning Pipeline for {LLM} in Retrieval-Augmented Generation

[16] Devlin et al., 2019, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Stay Ahead

Explore More Insights

Stay ahead with more perspectives on cutting-edge power, infrastructure, energy, bitcoin and AI solutions. Explore these articles to uncover strategies and insights shaping the future of industries.

AI-powered WAF, virtual patching: How F5 is hardening networks against frontier threats

“If the attacker is a machine and can devise new attack sequences in seconds, then your response to that cannot be signature-based. It has to be based around the behaviors that you detect and analyze,” Joel Moses, vice president of strategic engineering at F5, told Network World. Inside the AI-powered

A quick look at Cisco’s strategy to become a software monster

“What they are trying to do is get to a place where rather than just sell you a server or network switch and I’m done, is make themselves into basically a cloud service provider,” said Gold. At the core of Cisco’s strategy is its growing focus on security and network

Residential proxies are hiding in plain sight inside enterprise networks

Monthly query volume to those domains grew roughly 25% between January 2025 and April 2026, reaching over 500 billion queries per month. Residential proxy traffic appeared in every industry vertical examined, with at least 40% of customers in each sector affected. Over 90% of pharmaceutical and food and beverage customers

AI power efficiency the target of Lotus Microsystems energy advances

By shortening current paths and integrating thermal management directly into the power-delivery structure, vStrata aims to reduce conversion losses while improving cooling efficiency. According to Lotus Microsystems, the module can achieve point-of-load efficiencies of up to 96% while reducing power-conversion losses by more than 50% compared with conventional approaches. “We

Energy Department Issues RFP to Advance President Trump’s 172-Million-Barrel Strategic Petroleum Reserve Exchange

WASHINGTON—The U.S. Department of Energy (DOE) today issued a Request for Proposal (RFP) for an exchange of up to 40 million barrels of crude oil from the Strategic Petroleum Reserve (SPR). Today’s solicitation opens competitive bidding, continuing DOE’s execution of President Trump’s 172-million-barrel release as part of a coordinated 400-million-barrel action by International Energy Agency (IEA) member nations’ strategic reserves. Under President Trump’s leadership, DOE has advanced an unprecedented series of large-scale SPR exchange solicitations at record speed. These actions have moved critical crude oil supplies into the market to address short term supply disruptions and bolster energy security for the United States and its allies. The crude oil will originate from the SPR’s Big Hill and Bryan Mound sites. This action builds on the Department’s four previous solicitations that collectively awarded more than 133 million barrels across three completed exchanges. DOE’s earlier exchanges demonstrated the SPR’s ability to rapidly deliver crude under emergency authorities while achieving a 26 percent premium in returned barrels—expanding the reserve at no additional cost to American taxpayers. “With today’s announcement, we are accelerating the President’s commitment to a coordinated and strategic release that stabilizes global oil markets,” said DOE Acting Assistant Secretary for the Hydrocarbons and Geothermal Energy Office Curt Coccodrilli. “This exchange will help move oil swiftly to refiners, ease short-term supply pressures, and ensure the Strategic Petroleum Reserve continues to grow stronger through the return of premium barrels.” Under DOE’s exchange authority, participating companies will return the 40 million borrowed barrels with additional premium barrels, ensuring immediate market supply while increasing the SPR’s long-term inventory. Bids for this solicitation are due no later than 11:00 A.M. Central Time on Monday, June 15, 2026. For more information on the SPR, please visit DOE’s website.

DOE’s Hydrocarbons and Geothermal Energy Office Invests $3.6 Million to Modernize America’s Coal-Fired Power Plants

WASHINGTON—The U.S. Department of Energy’s (DOE) Hydrocarbons and Geothermal Energy Office (HGEO) today announced $3.6 million for nine design and engineering projects that will support the refurbishment or retrofit of existing coal power plants with transformational technologies that address wastewater systems and improve the efficiency, reliability, flexibility, and performance of coal and natural gas use. By upgrading our nation’s existing coal facilities, these initiatives will help strengthen the backbone of America’s power grid and ensure all American’s have access to affordable, reliable, and secure energy when they need it most. These efforts help to advance President Trump’s Executive Orders Reinvigorating America’s Beautiful Clean Coal Industry and Strengthening the Reliability and Security of the United States Electric Grid to restore common-sense energy policies that prioritize dependable power, affordability, and American workers. “America’s coal fleet is an undeniable pillar of our energy dominance and economic strength, but for too long, policies have undermined this vital industry and the dedicated workforce behind it, threatening our grid’s stability and driving up costs for everyday Americans,” said DOE Acting Assistant Secretary of the Hydrocarbons and Geothermal Energy Office Curt Coccodrilli. “With the project investments announced today, we are decisively moving to champion our existing coal plants, ensuring they continue to deliver affordable, reliable power, keep the lights on, and fuel America’s progress for generations to come.” Projects have been selected under three topic areas to provide a path forward to rapidly and cost-effectively restore the stability of the nation’s bulk power system while also finding beneficial uses for wastes generated by coal-based energy production. The projects will be executed in three phases, with design and engineering completed in Phase I, final engineering and detailed design completed in Phase II, and technology implementation and validation completed in Phase III. Selectees to receive Phase I funding include: Baker Hughes Energy Transition LLC (Houston, Texas),

Energy Department Releases Finalized Fusion Science and Technology Roadmap to Accelerate Commercial Fusion Power

WASHINGTON—The U.S. Department of Energy (DOE) today released the finalized Fusion Science and Technology (FS&T) Roadmap, a national strategy to accelerate the development and commercialization of fusion energy on the most rapid, responsible timeline in history. Building on earlier roadmap efforts, the finalized roadmap brings together fusion science, technology, infrastructure, workforce development, and commercialization priorities into a single national strategy to support fusion pilot plants and commercial fusion power in the mid-2030s. Fusion is the process that powers the sun and stars. For decades, scientists and engineers have worked to bring that same process to Earth as a source of abundant, reliable energy. The finalized roadmap outlines how DOE, industry, universities, and national laboratories will work together to accelerate the path toward commercial fusion energy in the United States. This effort advances President Trump’s energy dominance agenda and reinforces the Administration’s commitment to expanding reliable American energy production, strengthening domestic supply chains, and maintaining U.S. leadership in critical technologies. By accelerating progress toward commercial fusion power, DOE is helping secure a future of abundant and reliable energy. “Fusion energy has entered a new era defined by extraordinary scientific progress and public-private momentum,” said DOE Under Secretary for Science Dr. Darío Gil. “With this roadmap, we now have the clarity, coordination, and sustained commitment needed to turn the promise of fusion into a reality for the American people.” Developed with input from more than 800 scientists and engineers across the public and private sectors, the finalized FS&T Roadmap reflects contributions from more than 15 private companies, over 10 National Laboratories, and more than 70 universities. The roadmap identifies the critical science and technology gaps that must be closed to realize fusion pilot plants and strengthen U.S. leadership in the global fusion industry. The FS&T Roadmap establishes a unified strategy for the U.S.

Aramco to divest Malaysian refining assets

Petroliam Nasional Bhd. (PETRONAS) subsidiary PETRONAS Refinery & Petrochemical Corp. Sdn. Bhd. (PRPC) has agreed to buyout Saudi Arabian Oil Co.’s (Aramco) equity interests in the partners’ dual 50-50 joint ventures responsible for operating the 300,000-b/d integrated refining and petrochemical refinery of the Pengerang Integrated Complex (PIC) in southeastern Johor, Malaysia. Subject to fulfillment of customary closing conditions, Petronas will take 100% ownership and become full operator of Pengerang Refining Co. Sdn. Bhd. and Pengerang Petrochemical Co. Sdn. Bhd., collectively known as PRefChem, Aramco and Petronas said in separate releases. Aramco said divestment of the Malaysian assets will support the strategic optimization of the company’s own downstream portfolio by providing additional flexibility to pursue investments aligned with its broader downstream strategy. While Aramco will no longer hold ownership in the Malaysian ventures, the company said it will continue actively explore commercial arrangements with Petronas following the sale, including continuing its existing agreement to supply Saudi Arabian crude oil to the site, as well as opportunities related to technology exchange and integrated product distribution. Petronas said its move to take full control of the downstream assets will allow the company to further enhance operational alignment and flexibility across PRefChem’s value chain, while harnessing its international oil supply network and integrated operating model to support continued reliability and resilience across varying market conditions. Full ownership of PRefChem’s in-country operations also will strengthen Petronas’ ability to support Malaysia’s long-term energy security and industry resilience, the operator said. A definitive timeframe for when the parties expect to finalize the proposed transaction was not revealed. PRefChem operations In addition to the Johor refinery, PRefChem’s operations at PIC include a steam cracker complex equipped to produce 3.4 million tonnes/year (tpy) combined of ethylene, propylene, butadiene, benzene and raffinate-2. PRefChem also operates an associated petrochemical complex at the

Delfin Midstream takes $5-billion FID for first FLNG vessel

@import url(‘https://fonts.googleapis.com/css2?family=Inter:[email protected]&display=swap’); .ebm-page__main h1, .ebm-page__main h2, .ebm-page__main h3, .ebm-page__main h4, .ebm-page__main h5, .ebm-page__main h6 { font-family: Inter; } body { line-height: 150%; letter-spacing: 0.025em; } button, .ebm-button-wrapper { font-family: Inter; } .label-style { text-transform: uppercase; color: var(–color-grey); font-weight: 600; font-size: 0.75rem; } .caption-style { font-size: 0.75rem; opacity: .6; } #onetrust-pc-sdk [id*=btn-handler], #onetrust-pc-sdk [class*=btn-handler] { background-color: #c19a06 !important; border-color: #c19a06 !important; } #onetrust-policy a, #onetrust-pc-sdk a, #ot-pc-content a { color: #c19a06 !important; } #onetrust-consent-sdk #onetrust-pc-sdk .ot-active-menu { border-color: #c19a06 !important; } #onetrust-consent-sdk #onetrust-accept-btn-handler, #onetrust-banner-sdk #onetrust-reject-all-handler, #onetrust-consent-sdk #onetrust-pc-btn-handler.cookie-setting-link { background-color: #c19a06 !important; border-color: #c19a06 !important; } #onetrust-consent-sdk .onetrust-pc-btn-handler { color: #c19a06 !important; border-color: #c19a06 !important; } <!–> Delfin Midstream Inc., Houston, has taken a final investment decision (FID) for the first floating liquefied natural gas (FLNG) vessel of the Delfin LNG project under development in Louisiana and offshore in the Gulf of Mexico. Delfin FLNG 1 will be the first FLNG vessel in the US and the largest FLNG project globally, with an expected export capacity of 4.4 million tonnes/year (tpy) of LNG, the company said in its June 3 release. Concurrent with the FID, a group of investors led by Global Infrastructure Partners (GIP), a part of BlackRock—including existing Delfin investors Mitsui OSK Lines Ltd. (MOL), Vitol, and Diameter Capital Partners—has agreed to invest in the first phase of the project. ]–> <!–> –><!–> –> Oct. 9, 2023 <!–> –><!–> –> July 11, 2023 <!–> –><!–> –> June 9, 2023 <!–> –><!–> –> July 2, 2021 <!–> –> <!–> The vessel is backed by long-term LNG sales agreements with Vitol, Expand Energy, Centrica, and Gunvor, Delfin said, and all necessary permits and licenses required to begin construction have been secured. Construction contracts have been executed with Samsung Heavy Industries Co. Ltd. and Black & Veatch. LNG production is scheduled to begin

Chevron files $13.8-billion Argentina oil development proposal

Chevron Corp. applied June 2 to join Argentina’s Large Investment Incentive Regime (RIGI) for a $13.8-billion unconventional oil development at its 100% operated El Trapial-Este block in northern Neuquén province. Until recently, RIGI had attracted about $93 billion across 36 projects. Chevron’s application, which remains subject to government approval, is equivalent to almost one seventh of that total. The filing, which does not consitute a final investment decision, is Chevron’s largest individual investment proposal in Argentina since it entered the country in 1999 and the second-largest project submitted under RIGI, behind YPF SA’s $25-billion LLL Oil development. Chevron said it is targeting production of about 30,000 b/d from El Trapial-Este, subject to the availability of takeaway infrastructure. The block currently produces about 7,000 b/d. Chevron tested the block with a 7-well pilot in 2021 and has been carrying out development since late 2022, using laterals of more than 3,000 m and techniques transferred from the US Permian basin. In 2023, Chevron committed $500 million to that phase. During the company’s first-quarter earnings call on May 1, chief executive officer Mike Wirth anchored Chevron’s 2030 targets in “assets that are operating today.” El Trapial-Este was not explicitly identified among assets described as the main base for those targets. Wirth also said Chevron would not accelerate Permian production even with Brent above $100/bbl, preferring to manage that asset for free cash flow rather than volume. In the same presentation, Wirth named Argentina among the sources of equity crude that feed Chevron’s global refining system, along with Tengiz, Guyana, the Permian, and Venezuela. The earnings call came weeks before the El proposal filing. Vaca Muerta costs, takeaway capacity Breakeven costs in Vaca Muerta’s best blocks are about $40/bbl at the wellhead, according to Rystad Energy, while normalized well productivity—adjusted for lateral length and fracture

From the data center to the edge: How to build secure, effective enterprise AI infrastructure

While hyperscalers and neo-cloud providers may get the lion’s share of attention for providing AI infrastructure, many enterprises are taking a build-it-themselves approach to meet their specific AI requirements. The success of such projects is crucial to achieving business objectives, yet companies face significant challenges as they try to scale pilots to production. Organizations must keep up with the dynamic, ever-changing demands that AI applications place on compute and network infrastructure, from the data center to the edge. That means architecting systems to grow as demand warrants and to avoid performance bottlenecks. The architecture must also account for AI-driven security vulnerabilities and ensure appropriate defenses are in place. Yes, it’s a tall order. But here, in simplified form, is a three-step plan for meeting those objectives. Step one: Go modular Integrating all the required components in piecemeal fashion for an AI factory is complex, costly, and fraught with integration risk. Start with a modular design, based on proven NVIDIA reference architectures. A modular approach combines pre-validated accelerated computing hardware, AI software, and orchestration platforms, as well as networking and storage capabilities. A modular strategy speeds implementation and creates a faster time to value for your AI infrastructure. Using modules that combine compute, networking, and storage makes it easier to scale capacity as needed, whether in the data center or at edge facilities. In addition, the modular approach simplifies the job of addressing varying requirements, from inferencing engines at the edge to massive-scale model training in the data center, while staying within the same solution family. The same applies to easing integration processes, as modular platforms offer pre-validated software. The Cisco Secure AI Factory with NVIDIA approach, for example, includes hardware (Cisco AI PODS) that is pre-validated to work with NVIDIA AI Enterprise software; Cisco Security and Splunk Observability software; orchestration

OpenAI weighs Nvidia-backed lease for 10 GW Ohio data center campus

OpenAI would control the computing equipment under a 20-year lease and begin payments once the site starts operating, with the first phase expected in 2028. Nvidia is expected to supply the hardware and guarantee both OpenAI’s lease obligations and the developer’s financing, the report added. The reported structure highlights a broader shift in AI infrastructure strategy, where model developers, chip suppliers, and energy providers are forging increasingly long-term partnerships to secure compute capacity amid surging demand. “These types of symbiotic deals are becoming the norm as AI infrastructure rolls out,” said Neil Shah, vice president for research and partner at Counterpoint Research. “If a CIO picks OpenAI to be the base layer, they shouldn’t just accept whatever infrastructure comes with it. CIOs need to negotiate and demand that OpenAI uses a mix of capacity so all your eggs are not in one premium basket like Nvidia.” OpenAI and Nvidia did not immediately respond to requests for comment.

Arista unveils 1.6T rack-scale switch family for AI infrastructure

The new Arista family joins a growing ecosystem of vendors looking to tap into the 1.6T Ethernet world, which includes Cisco, Nvidia, Celestica and others. “Arista Network’s new 7060XE7 Series is a strong signal of where large-scale AI fabrics are heading: higher bandwidth, better power efficiency, and tighter integration between compute, optics, silicon, cooling, and network operating software,” wrote Sameh Boujelbene, vice president, data center switch and AI networks market research for Dell Oro, in a LinkedIn post. Among the features that stand out to her are “strong customer and ecosystem validation from Microsoft Azure, Oracle Cloud Infrastructure, Meta, AMD, and Broadcom.”

Water Emerges as a Critical Constraint for AI Data Centers

“There really has been a major shift within the last couple of years,” Bajpayee said. “I would even say within the last 12 months is where we have seen suddenly a rapid increase in the data center operators’ desire to control their water destiny.” For Gradiant, the MIT-born water technology company that built its reputation serving semiconductor manufacturers, pharmaceutical companies, and industrial customers worldwide, that shift has translated into a rapidly expanding pipeline of data center opportunities. More importantly, Bajpayee believes it signals a fundamental change in how the industry thinks about water itself. The conversation is no longer centered primarily on sustainability metrics or corporate environmental goals. Instead, operators increasingly view water as a business continuity issue. “We’re seeing operators themselves come to us and tell us that these are issues they are facing,” Bajpayee said. “They want to make sure they don’t get stalled, their permits don’t get pulled, their business doesn’t get stopped, and communities don’t push them out because they didn’t figure out a way to control their water.” From Water Treatment to Water Strategy That shift is occurring as Gradiant expands deployments of its recently announced HyperSolved platform, an end-to-end cooling water management system purpose-built for AI data centers. The company says HyperSolved is now being deployed with several of the world’s largest hyperscale operators across North America, Europe, and Asia, reflecting growing industry demand for integrated approaches to water infrastructure. While compute, networking, and power systems have evolved rapidly during the AI era, water management often remains fragmented, requiring operators to coordinate multiple vendors responsible for sourcing, treatment, cooling, wastewater management, reuse, discharge, and regulatory compliance. Gradiant’s approach seeks to consolidate those functions into a single integrated platform and operating model. The timing reflects the growing scale of the challenge. New AI data center

Data Center Jobs: Engineering, Construction, Commissioning, Sales, Field Service and Facility Tech Jobs Available in Major Data Center Hotspots

Each month Data Center Frontier, in partnership with Pkaza, posts some of the hottest data center career opportunities in the market. Here’s a look at some of the latest data center jobs posted on the Data Center Frontier jobs board, powered by Pkaza Critical Facilities Recruiting. Looking for Data Center Candidates? Check out Pkaza’s Active Candidate / Featured Candidate Hotlist Mechanical Applications Engineer Pittsburgh, PA This position is also available in: Denver, CO; Richmond, VA and Georgetown, SC (live by the beach!). Relo available. Our client is a leading provider and manufacturer of industrial HVAC mechanical equipment used in industrial cooling applications for mission critical operations. They help their customers save money by reducing energy and operating costs and provide solutions for modernizing their customer’s existing mechanical infrastructure. This company provides cooling solutions to many of the world’s largest organizations and government facilities and enterprise clients, colocation providers and hyperscale companies. This career-growth minded opportunity offers exciting projects with leading-edge technology and innovation as well as competitive salaries and benefits. Electrical Commissioning Engineer New Albany, OH This traveling position is also available in: New York, NY; White Plains, NY; Dallas, TX; Richmond, VA; Ashburn, VA; Montvale, NJ; Charlotte, NC; Atlanta, GA; Hampton, GA; Cedar Rapids, IA; Phoenix, AZ; Salt Lake City, UT; Kansas City, MO; Omaha, NE; Chesterton, IN; Indianapolis, IN or Chicago, IL. *** ALSO looking for a LEAD EE and ME CxA Agents and CxA PMs *** Our client is an engineering design and commissioning company that has a national footprint and specializes in MEP critical facilities design. They provide design, commissioning, consulting and management expertise in the critical facilities space. They have a mindset to provide reliability, energy efficiency, sustainable design and LEED expertise when providing these consulting services for Enterprise, Colocation and Hyperscale Companies. This career-growth minded opportunity offers exciting projects

Fiber’s Next Act: How AI Is Driving Connectivity Closer to the Edge

ORLANDO, Fla. — Much of the conversation surrounding AI infrastructure has focused on GPUs, power generation, cooling systems, and the unprecedented scale of next-generation data center development. But at Fiber Connect 2026, another reality became increasingly clear: none of those investments matter without the network infrastructure required to connect them. That theme emerged repeatedly during a conversation between Data Center Frontier Editor-in-Chief Matt Vincent and Clearfield Chief Commercial Officer Anis Khemakhem, whose perspective sits at the intersection of broadband infrastructure, fiber deployment, and emerging AI connectivity requirements. While Clearfield is best known throughout the broadband industry for its fiber management and connectivity solutions, Khemakhem argued that AI’s rapid expansion is creating new opportunities, and new challenges, that extend well beyond traditional fiber-to-the-home deployments. “AI is driving that connectivity closer and closer to the edge,” Khemakhem said, noting that growing compute requirements and increasingly latency-sensitive workloads are fundamentally changing assumptions about where infrastructure must reside and how it must be connected. For Data Center Frontier readers, the significance lies in a growing realization that AI infrastructure is becoming as much a networking challenge as a compute challenge. Beyond the Traditional Data Center One of the more notable themes of the discussion was Khemakhem’s view that the term “data center” has become too broad to be useful. The industry often speaks of data centers as a single category, but Clearfield increasingly differentiates between hyperscale campuses, colocation facilities, central office environments, and a rapidly emerging class of edge deployments. “There is no one-size-fits-all data center,” Khemakhem said, describing a continuum that extends from hyperscale facilities all the way to edge locations positioned near users and applications. That distinction matters because many AI applications are introducing latency requirements that cannot always be addressed by centralized facilities alone. As AI inference moves closer to users,

Microsoft will invest $80B in AI data centers in fiscal 2025

And Microsoft isn’t the only one that is ramping up its investments into AI-enabled data centers. Rival cloud service providers are all investing in either upgrading or opening new data centers to capture a larger chunk of business from developers and users of large language models (LLMs). In a report published in October 2024, Bloomberg Intelligence estimated that demand for generative AI would push Microsoft, AWS, Google, Oracle, Meta, and Apple would between them devote $200 billion to capex in 2025, up from $110 billion in 2023. Microsoft is one of the biggest spenders, followed closely by Google and AWS, Bloomberg Intelligence said. Its estimate of Microsoft’s capital spending on AI, at $62.4 billion for calendar 2025, is lower than Smith’s claim that the company will invest $80 billion in the fiscal year to June 30, 2025. Both figures, though, are way higher than Microsoft’s 2020 capital expenditure of “just” $17.6 billion. The majority of the increased spending is tied to cloud services and the expansion of AI infrastructure needed to provide compute capacity for OpenAI workloads. Separately, last October Amazon CEO Andy Jassy said his company planned total capex spend of $75 billion in 2024 and even more in 2025, with much of it going to AWS, its cloud computing division.

John Deere unveils more autonomous farm machines to address skill labor shortage

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Self-driving tractors might be the path to self-driving cars. John Deere has revealed a new line of autonomous machines and tech across agriculture, construction and commercial landscaping. The Moline, Illinois-based John Deere has been in business for 187 years, yet it’s been a regular as a non-tech company showing off technology at the big tech trade show in Las Vegas and is back at CES 2025 with more autonomous tractors and other vehicles. This is not something we usually cover, but John Deere has a lot of data that is interesting in the big picture of tech. The message from the company is that there aren’t enough skilled farm laborers to do the work that its customers need. It’s been a challenge for most of the last two decades, said Jahmy Hindman, CTO at John Deere, in a briefing. Much of the tech will come this fall and after that. He noted that the average farmer in the U.S. is over 58 and works 12 to 18 hours a day to grow food for us. And he said the American Farm Bureau Federation estimates there are roughly 2.4 million farm jobs that need to be filled annually; and the agricultural work force continues to shrink. (This is my hint to the anti-immigration crowd). John Deere’s autonomous 9RX Tractor. Farmers can oversee it using an app. While each of these industries experiences their own set of challenges, a commonality across all is skilled labor availability. In construction, about 80% percent of contractors struggle to find skilled labor. And in commercial landscaping, 86% of landscaping business owners can’t find labor to fill open positions, he said. “They have to figure out how to do

2025 playbook for enterprise AI success, from agents to evals

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More 2025 is poised to be a pivotal year for enterprise AI. The past year has seen rapid innovation, and this year will see the same. This has made it more critical than ever to revisit your AI strategy to stay competitive and create value for your customers. From scaling AI agents to optimizing costs, here are the five critical areas enterprises should prioritize for their AI strategy this year. 1. Agents: the next generation of automation AI agents are no longer theoretical. In 2025, they’re indispensable tools for enterprises looking to streamline operations and enhance customer interactions. Unlike traditional software, agents powered by large language models (LLMs) can make nuanced decisions, navigate complex multi-step tasks, and integrate seamlessly with tools and APIs. At the start of 2024, agents were not ready for prime time, making frustrating mistakes like hallucinating URLs. They started getting better as frontier large language models themselves improved. “Let me put it this way,” said Sam Witteveen, cofounder of Red Dragon, a company that develops agents for companies, and that recently reviewed the 48 agents it built last year. “Interestingly, the ones that we built at the start of the year, a lot of those worked way better at the end of the year just because the models got better.” Witteveen shared this in the video podcast we filmed to discuss these five big trends in detail. Models are getting better and hallucinating less, and they’re also being trained to do agentic tasks. Another feature that the model providers are researching is a way to use the LLM as a judge, and as models get cheaper (something we’ll cover below), companies can use three or more models to

OpenAI’s red teaming innovations define new essentials for security leaders in the AI era

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More OpenAI has taken a more aggressive approach to red teaming than its AI competitors, demonstrating its security teams’ advanced capabilities in two areas: multi-step reinforcement and external red teaming. OpenAI recently released two papers that set a new competitive standard for improving the quality, reliability and safety of AI models in these two techniques and more. The first paper, “OpenAI’s Approach to External Red Teaming for AI Models and Systems,” reports that specialized teams outside the company have proven effective in uncovering vulnerabilities that might otherwise have made it into a released model because in-house testing techniques may have missed them. In the second paper, “Diverse and Effective Red Teaming with Auto-Generated Rewards and Multi-Step Reinforcement Learning,” OpenAI introduces an automated framework that relies on iterative reinforcement learning to generate a broad spectrum of novel, wide-ranging attacks. Going all-in on red teaming pays practical, competitive dividends It’s encouraging to see competitive intensity in red teaming growing among AI companies. When Anthropic released its AI red team guidelines in June of last year, it joined AI providers including Google, Microsoft, Nvidia, OpenAI, and even the U.S.’s National Institute of Standards and Technology (NIST), which all had released red teaming frameworks. Investing heavily in red teaming yields tangible benefits for security leaders in any organization. OpenAI’s paper on external red teaming provides a detailed analysis of how the company strives to create specialized external teams that include cybersecurity and subject matter experts. The goal is to see if knowledgeable external teams can defeat models’ security perimeters and find gaps in their security, biases and controls that prompt-based testing couldn’t find. What makes OpenAI’s recent papers noteworthy is how well they define using human-in-the-middle