Stay Ahead, Stay ONMINE

Mastering the Poisson Distribution: Intuition and Foundations

You’ve probably used the normal distribution one or two times too many. We all have — It’s a true workhorse. But sometimes, we run into problems. For instance, when predicting or forecasting values, simulating data given a particular data-generating process, or when we try to visualise model output and explain them intuitively to non-technical stakeholders. Suddenly, things don’t make much sense: can a user really have made -8 clicks on the banner? Or even 4.3 clicks? Both are examples of how count data doesn’t behave. I’ve found that better encapsulating the data generating process into my modelling has been key to having sensible model output. Using the Poisson distribution when it was appropriate has not only helped me convey more meaningful insights to stakeholders, but it has also enabled me to produce more accurate error estimates, better Inference, and sound decision-making. In this post, my aim is to help you get a deep intuitive feel for the Poisson distribution by walking through example applications, and taking a dive into the foundations — the maths. I hope you learn not just how it works, but also why it works, and when to apply the distribution. If you know of a resource that has helped you grasp the concepts in this blog particularly well, you’re invited to share it in the comments! Outline Examples and use cases: Let’s walk through some use cases and sharpen the intuition I just mentioned. Along the way, the relevance of the Poisson Distribution will become clear. The foundations: Next, let’s break down the equation into its individual components. By studying each part, we’ll uncover why the distribution works the way it does. The assumptions: Equipped with some formality, it will be easier to understand the assumptions that power the distribution, and at the same time set the boundaries for when it works, and when not. When real life deviates from the model: Finally, let’s explore the special links that the Poisson distribution has with the Negative Binomial distribution. Understanding these relationships can deepen our understanding, and provide alternatives when the Poisson distribution is not suited for the job. Example in an online marketplace I chose to deep dive into the Poisson distribution because it frequently appears in my day-to-day work. Online marketplaces rely on binary user choices from two sides: a seller deciding to list an item and a buyer deciding to make a purchase. These micro-behaviours drive supply and demand, both in the short and long term. A marketplace is born. Binary choices aggregate into counts — the sum of many such decisions as they occur. Attach a timeframe to this counting process, and you’ll start seeing Poisson distributions everywhere. Let’s explore a concrete example next. Consider a seller on a platform. In a given month, the seller may or may not list an item for sale (a binary choice). We would only know if she did because then we’d have a measurable count of the event. Nothing stops her from listing another item in the same month. If she does, we count those events. The total could be zero for an inactive seller or, say, 120 for a highly engaged seller. Over several months, we would observe a varying number of listed items by this seller — sometimes fewer, sometimes more — hovering around an average monthly listing rate. That is essentially a Poisson process. When we get to the assumptions section, you’ll see what we had to assume away to make this example work. Other examples Other phenomena that can be modelled with a Poisson distribution include: Sports analytics: The number of goals scored in a match between two teams. Queuing: Customers arriving at a help desk or customer support calls. Insurance: The number of claims made within a given period. Each of these examples warrants further inspection, but for the remainder of this post, we’ll use the marketplace example to illustrate the inner workings of the distribution. The mathy bit … or foundations. I find opening up the probability mass function (PMF) of distributions helpful to understanding why things work as they do. The PMF of the Poisson distribution goes like: Where λ is the rate parameter, and 𝑘 is the manifested count of the random variable (𝑘 = 0, 1, 2, 3, … events). Very neat and compact. The probability mass function of the Poisson distribution, for a few different lambdas. Contextualising λ and k: the marketplace example In the context of our earlier example — a seller listing items on our platform — λ represents the seller’s average monthly listings. As the expected monthly value for this seller, λ orchestrates the number of items she would list in a month. Note that λ is a Greek letter, so read: λ is a parameter that we can estimate from data. On the other hand, 𝑘 does not hold any information about the seller’s idiosyncratic behaviour. It’s the target value we set for the number of events that may happen to learn about its probability. The dual role of λ as the mean and variance When I said that λ orchestrates the number of monthly listings for the seller, I meant it quite literally. Namely, λ is both the expected value and variance of the distribution, indifferently, for all values of λ. This means that the mean-to-variance ratio (index of dispersion) is always 1. To put this into perspective, the normal distribution requires two parameters — 𝜇 and 𝜎², the average and variance respectively — to fully describe it. The Poisson distribution achieves the same with just one. Having to estimate only one parameter can be beneficial for parametric inference. Specifically, by reducing the variance of the model and increasing the statistical power. On the other hand, it can be too limiting of an assumption. Alternatives like the Negative Binomial distribution can alleviate this limitation. We’ll explore that later. Breaking down the probability mass function Now that we know the smallest building blocks, let’s zoom out one step: what is λᵏ, 𝑒^⁻λ, and 𝑘!, and more importantly, what is each of these components’ function in the whole? λᵏ is a weight that expresses how likely it is for 𝑘 events to happen, given that the expectation is λ. Note that “likely” here does not mean a probability, yet. It’s merely a signal strength. 𝑘! is a combinatorial correction so that we can say that the order of the events is irrelevant. The events are interchangeable. 𝑒^⁻λ normalises the integral of the PMF function to sum up to 1. It’s called the partition function of exponential-family distributions. In more detail, λᵏ relates the observed value 𝑘 to the expected value of the random variable, λ. Intuitively, more probability mass lies around the expected value. Hence, if the observed value lies close to the expectation, the probability of occurring is larger than the probability of an observation far removed from the expectation. Before we can cross-check our intuition with the numerical behaviour of λᵏ, we need to consider what 𝑘! does. Interchangeable events Had we cared about the order of events, then each unique event could be ordered in 𝑘! ways. But because we don’t, and we deem each event interchangeable, we “divide out” 𝑘! from λᵏ to correct for the overcounting. Since λᵏ is an exponential term, the output will always be larger as 𝑘 grows, holding λ constant. That is the opposite of our intuition that there is maximum probability when λ = 𝑘, as the output is larger when 𝑘 = λ + 1. But now that we know about the interchangeable events assumption — and the overcounting issue — we know that we have to factor in 𝑘! like so: λᵏ 𝑒^⁻λ / 𝑘!, to see the behaviour we expect. Now let’s check the intuition of the relationship between λ and 𝑘 through λᵏ, corrected for 𝑘!. For the same λ, say λ = 4, we should see λᵏ 𝑒^⁻λ / 𝑘! to be smaller for values of 𝑘 that are far removed from 4, compared to values of 𝑘 that lie close to 4. Like so: inline code: 4²/2 = 8 is smaller than 4⁴/24 = 10.7. This is consistent with the intuition of a higher likelihood of 𝑘 when it’s near the expectation. The image below shows this relationship more generally, where you see that the output is larger as 𝑘 approaches λ. The probability mass function without the normalising component e^-lambda. The assumptions First, let’s get one thing off the table: the difference between a Poisson process, and the Poisson distribution. The process is a stochastic continuous-time model of points happening in given interval: 1D, a line; 2D, an area, or higher dimensions. We, data scientists, most often deal with the one-dimensional case, where the “line” is time, and the points are the events of interest — I dare to say. These are the assumptions of the Poisson process: The occurrence of one event does not affect the probability of a second event. Think of our seller going on to list another item tomorrow indifferently of having done so already today, or the one from five days ago for that matter. The point here is that there is no memory between events. The average rate at which events occur, is independent of any occurrence. In other words, no event that happened (or will happen) alters λ, which remains constant throughout the observed timeframe. In our seller example, this means that listing an item today does not increase or decrease the seller’s motivation or likelihood of listing another item tomorrow. Two events cannot occur at exactly the same instant. If we were to zoom at an infinite granular level on the timescale, no two listings could have been placed simultaneously; always sequentially. From these assumptions — no memory, constant rate, events happening alone — it follows that 1) any interval’s number of events is Poisson-distributed with parameter λₜ and 2) that disjoint intervals are independent — two key properties of a Poisson process. A Note on the distribution:The distribution simply describes probabilities for various numbers of counts in an interval. Strictly speaking, one can use the distribution pragmatically whenever the data is nonnegative, can be unbounded on the right, has mean λ, and reasonably models the data. It would be just convenient if the underlying process is a Poisson one, and actually justifies using the distribution. The marketplace example: Implications So, can we justify using the Poisson distribution for our marketplace example? Let’s open up the assumptions of a Poisson process and take the test. Constant λ Why it may fail: The seller has patterned online activity; holidays; promotions; listings are seasonal goods. Consequence: λ is not constant, leading to overdispersion (mean-to-variance ratio is larger than 1, or to temporal patterns. Independence and memorylessness Why it may fail: The propensity to list again is higher after a successful listing, or conversely, listing once depletes the stock and intervenes with the propensity of listing again. Consequence: Two events are no longer independent, as the occurrence of one informs the occurrence of the other. Simultaneous events Why it may fail: Batch-listing, a new feature, was introduced to help the sellers. Consequence: Multiple listings would come online at the same time, clumped together, and they would be counted simultaneously. Balancing rigour and pragmatism As Data Scientists on the job, we may feel trapped between rigour and pragmatism. The three steps below should give you a sound foundation to decide on which side to err, when the Poisson distribution falls short: Pinpoint your goal: is it inference, simulation or prediction, and is it about high-stakes output? List the worst thing that can happen, and the cost of it for the business. Identify the problem and solution: why does the Poisson distribution not fit, and what can you do about it? list 2-3 solutions, including changing nothing. Balance gains and costs: Will your workaround improve things, or make it worse? and at what cost: interpretability, new assumptions introduced and resources used. Does it help you in achieving your goal? That said, here are some counters I use when needed. When real life deviates from your model Everything described so far pertains to the standard, or homogenous, Poisson process. But what if reality begs for something different? In the next section, we’ll cover two extensions of the Poisson distribution when the constant λ assumption does not hold. These are not mutually exclusive, but neither they are the same: Time-varying λ: a single seller whose listing rate ramps up before holidays and slows down afterward Mixed Poisson distribution: multiple sellers listing items, each with their own λ can be seen as a mixture of various Poisson processes Time-varying λ The first extension allows λ to have its own value for each time t. The PMF then becomes Where the number of events 𝐾(𝑇) in an interval 𝑇 follows the Poisson distribution with a rate no longer equal to a fixed λ, but one equal to: More intuitively, integrating over the interval 𝑡 to 𝑡 + 𝑖 gives us a single number: the expected value of events over that interval. The integral will vary by each arbitrary interval, and that’s what makes λ change over time. To understand how that integration works, it was helpful for me to think of it like this: if the interval 𝑡 to 𝑡₁ integrates to 3, and 𝑡₁ to 𝑡₂ integrates to 5, then the interval 𝑡 to 𝑡₂ integrates to 8 = 3 + 5. That’s the two expectations summed up, and now the expectation of the entire interval. Practical implication One may want to modeling the expected value of the Poisson distribution as a function of time. For instance, to model an overall change in trend, or seasonality. In generative model notation: Time may be a continuous variable, or an arbitrary function of it. Process-varying λ: Mixed Poisson distribution But then there’s a gotcha. Remember when I said that λ has a dual role as the mean and variance? That still applies here. Looking at the “relaxed” PMF*, the only thing that changes is that λ can vary freely with time. But it’s still the one and only λ that orchestrates both the expected value and the dispersion of the PMF*. More precisely, 𝔼[𝑋] = Var(𝑋) still holds. There are various reasons for this constraint not to hold in reality. Model misspecification, event interdependence and unaccounted for heterogeneity could be the issues at hand. I’d like to focus on the latter case, as it justifies the Negative Binomial distribution — one of the topics I promised to open up. Heterogeneity and overdispersionImagine we are not dealing with one seller, but with 10 of them listing at different intensity levels, λᵢ, where 𝑖 = 1, 2, 3, …, 10 sellers. Then, essentially, we have 10 Poisson processes going on. If we unify the processes and estimate the grand λ, we simplify the mixture away. Meaning, we get a correct estimate of all sellers on average, but the resulting grand λ is naive and does not know about the original spread of λᵢ. It still assumes that the variance and mean are equal, as per the axioms of the distribution. This will lead to overdispersion and, in turn, to underestimated errors. Ultimately, it inflates the false positive rate and drives poor decision-making. We need a way to embrace the heterogeneity amongst sellers’ λᵢ. Negative binomial: Extending the Poisson distributionAmong the few ways one can look at the Negative Binomial distribution, one way is to see it as a compound Poisson process — 10 sellers, sounds familiar yet? That means multiple independent Poisson processes are summed up to a single one. Mathematically, first we draw λ from a Gamma distribution: λ ~ Γ(r, θ), then we draw the count 𝑋 | λ ~ Poisson(λ). In one image, it is as if we would sample from plenty Poisson distributions, corresponding to each seller. A negative Binomial distribution arises from many Poisson distributions. The more exposing alias of the Negative binomial distribution is Gamma-Poisson mixture distribution, and now we know why: the dictating λ comes from a continuous mixture. That’s what we needed to explain the heterogeneity amongst sellers. Let’s simulate this scenario to gain more intuition. Gamma mixture of lambda. First, we draw λᵢ from a Gamma distribution: λᵢ ~ Γ(r, θ). Intuitively, the Gamma distribution tells us about the variety in the intensity — listing rate — amongst the sellers. On a practical note, one can instill their assumptions about the degree of heterogeneity in this step of the model: how different are sellers? By varying the levels of heterogeneity, one can observe the impact on the final Poisson-like distribution. Doing this type of checks (i.e., posterior predictive check), is common in Bayesian modeling, where the assumptions are set explicitly. Gamma-Poisson mixture distribution versus homogenous Poisson distribution. Τhe dashed line reflects λ, which is 4 for both distributions. In the second step, we plug the obtained λ into the Poisson distribution: 𝑋 | λ ~ Poisson(λ), and obtain a Poisson-like distribution that represents the summed subprocesses. Notably, this unified process has a larger dispersion than expected from a homogeneous Poisson distribution, but it is in line with the Gamma mixture of λ. Heterogeneous λ and inference A practical consequence of introducing flexibility into your assumed distribution is that inference becomes more challenging. More parameters (i.e., the Gamma parameters) need to be estimated. Parameters act as flexible explainers of the data, tending to overfit and explain away variance in your variable. The more parameters you have, the better the explanation may seem, but the model also becomes more susceptible to noise in the data. Higher variance reduces the power to identify a difference in means, if one exists, because — well — it gets lost in the variance. Countering the loss of power Confirm whether you indeed need to extend the standard Poisson distribution. If not, simplify to the simplest, most fit model. A quick check on overdispersion may suffice for this. Pin down the estimates of the Gamma mixture distribution parameters using regularising, informative priors (think: Bayes). During my research process for writing this blog, I learned a great deal about the connective tissue underlying all of this: how the binomial distribution plays a fundamental role in the processes we’ve discussed. And while I’d love to ramble on about this, I’ll save it for another post, perhaps. In the meantime, feel free to share your understanding in the comments section below 👍. Conclusion The Poisson distribution is a simple distribution that can be highly suitable for modelling count data. However, when the assumptions do not hold, one can extend the distribution by allowing the rate parameter to vary as a function of time or other factors, or by assuming subprocesses that collectively make up the count data. This added flexibility can address the limitations, but it comes at a cost: increased flexibility in your modelling raises the variance and, consequently, undermines the statistical power of your model. If your end goal is inference, you may want to think twice and consider exploring simpler models for the data. Alternatively, switch to the Bayesian paradigm and leverage its built-in solution to regularise estimates: informative priors. I hope this has given you what you came for — a better intuition about the Poisson distribution. I’d love to hear your thoughts about this in the comments! Unless otherwise noted, all images are by the author.Originally published at https://aalvarezperez.github.io on January 5, 2025.

You’ve probably used the normal distribution one or two times too many. We all have — It’s a true workhorse. But sometimes, we run into problems. For instance, when predicting or forecasting values, simulating data given a particular data-generating process, or when we try to visualise model output and explain them intuitively to non-technical stakeholders. Suddenly, things don’t make much sense: can a user really have made -8 clicks on the banner? Or even 4.3 clicks? Both are examples of how count data doesn’t behave.

I’ve found that better encapsulating the data generating process into my modelling has been key to having sensible model output. Using the Poisson distribution when it was appropriate has not only helped me convey more meaningful insights to stakeholders, but it has also enabled me to produce more accurate error estimates, better Inference, and sound decision-making.

In this post, my aim is to help you get a deep intuitive feel for the Poisson distribution by walking through example applications, and taking a dive into the foundations — the maths. I hope you learn not just how it works, but also why it works, and when to apply the distribution.

If you know of a resource that has helped you grasp the concepts in this blog particularly well, you’re invited to share it in the comments!

Outline

  1. Examples and use cases: Let’s walk through some use cases and sharpen the intuition I just mentioned. Along the way, the relevance of the Poisson Distribution will become clear.
  2. The foundations: Next, let’s break down the equation into its individual components. By studying each part, we’ll uncover why the distribution works the way it does.
  3. The assumptions: Equipped with some formality, it will be easier to understand the assumptions that power the distribution, and at the same time set the boundaries for when it works, and when not.
  4. When real life deviates from the model: Finally, let’s explore the special links that the Poisson distribution has with the Negative Binomial distribution. Understanding these relationships can deepen our understanding, and provide alternatives when the Poisson distribution is not suited for the job.

Example in an online marketplace

I chose to deep dive into the Poisson distribution because it frequently appears in my day-to-day work. Online marketplaces rely on binary user choices from two sides: a seller deciding to list an item and a buyer deciding to make a purchase. These micro-behaviours drive supply and demand, both in the short and long term. A marketplace is born.

Binary choices aggregate into counts — the sum of many such decisions as they occur. Attach a timeframe to this counting process, and you’ll start seeing Poisson distributions everywhere. Let’s explore a concrete example next.

Consider a seller on a platform. In a given month, the seller may or may not list an item for sale (a binary choice). We would only know if she did because then we’d have a measurable count of the event. Nothing stops her from listing another item in the same month. If she does, we count those events. The total could be zero for an inactive seller or, say, 120 for a highly engaged seller.

Over several months, we would observe a varying number of listed items by this seller — sometimes fewer, sometimes more — hovering around an average monthly listing rate. That is essentially a Poisson process. When we get to the assumptions section, you’ll see what we had to assume away to make this example work.

Other examples

Other phenomena that can be modelled with a Poisson distribution include:

  • Sports analytics: The number of goals scored in a match between two teams.
  • Queuing: Customers arriving at a help desk or customer support calls.
  • Insurance: The number of claims made within a given period.

Each of these examples warrants further inspection, but for the remainder of this post, we’ll use the marketplace example to illustrate the inner workings of the distribution.

The mathy bit

… or foundations.

I find opening up the probability mass function (PMF) of distributions helpful to understanding why things work as they do. The PMF of the Poisson distribution goes like:

Where λ is the rate parameter, and 𝑘 is the manifested count of the random variable (𝑘 = 0, 1, 2, 3, … events). Very neat and compact.

Graph: The probability mass function of the Poisson distribution, for a few different lambdas.
The probability mass function of the Poisson distribution, for a few different lambdas.

Contextualising λ and k: the marketplace example

In the context of our earlier example — a seller listing items on our platform — λ represents the seller’s average monthly listings. As the expected monthly value for this seller, λ orchestrates the number of items she would list in a month. Note that λ is a Greek letter, so read: λ is a parameter that we can estimate from data. On the other hand, 𝑘 does not hold any information about the seller’s idiosyncratic behaviour. It’s the target value we set for the number of events that may happen to learn about its probability.

The dual role of λ as the mean and variance

When I said that λ orchestrates the number of monthly listings for the seller, I meant it quite literally. Namely, λ is both the expected value and variance of the distribution, indifferently, for all values of λ. This means that the mean-to-variance ratio (index of dispersion) is always 1.

To put this into perspective, the normal distribution requires two parameters — 𝜇 and 𝜎², the average and variance respectively — to fully describe it. The Poisson distribution achieves the same with just one.

Having to estimate only one parameter can be beneficial for parametric inference. Specifically, by reducing the variance of the model and increasing the statistical power. On the other hand, it can be too limiting of an assumption. Alternatives like the Negative Binomial distribution can alleviate this limitation. We’ll explore that later.

Breaking down the probability mass function

Now that we know the smallest building blocks, let’s zoom out one step: what is λᵏ, 𝑒^⁻λ, and 𝑘!, and more importantly, what is each of these components’ function in the whole?

  • λᵏ is a weight that expresses how likely it is for 𝑘 events to happen, given that the expectation is λ. Note that “likely” here does not mean a probability, yet. It’s merely a signal strength.
  • 𝑘! is a combinatorial correction so that we can say that the order of the events is irrelevant. The events are interchangeable.
  • 𝑒^⁻λ normalises the integral of the PMF function to sum up to 1. It’s called the partition function of exponential-family distributions.

In more detail, λᵏ relates the observed value 𝑘 to the expected value of the random variable, λ. Intuitively, more probability mass lies around the expected value. Hence, if the observed value lies close to the expectation, the probability of occurring is larger than the probability of an observation far removed from the expectation. Before we can cross-check our intuition with the numerical behaviour of λᵏ, we need to consider what 𝑘! does.

Interchangeable events

Had we cared about the order of events, then each unique event could be ordered in 𝑘! ways. But because we don’t, and we deem each event interchangeable, we “divide out” 𝑘! from λᵏ to correct for the overcounting.

Since λᵏ is an exponential term, the output will always be larger as 𝑘 grows, holding λ constant. That is the opposite of our intuition that there is maximum probability when λ = 𝑘, as the output is larger when 𝑘 = λ + 1. But now that we know about the interchangeable events assumption — and the overcounting issue — we know that we have to factor in 𝑘! like so: λᵏ 𝑒^⁻λ / 𝑘!, to see the behaviour we expect.

Now let’s check the intuition of the relationship between λ and 𝑘 through λᵏ, corrected for 𝑘!. For the same λ, say λ = 4, we should see λᵏ 𝑒^⁻λ / 𝑘! to be smaller for values of 𝑘 that are far removed from 4, compared to values of 𝑘 that lie close to 4. Like so: inline code: 4²/2 = 8 is smaller than 4⁴/24 = 10.7. This is consistent with the intuition of a higher likelihood of 𝑘 when it’s near the expectation. The image below shows this relationship more generally, where you see that the output is larger as 𝑘 approaches λ.

Graph: The probability mass function without the normalising component e^-lambda.
The probability mass function without the normalising component e^-lambda.

The assumptions

First, let’s get one thing off the table: the difference between a Poisson process, and the Poisson distribution. The process is a stochastic continuous-time model of points happening in given interval: 1D, a line; 2D, an area, or higher dimensions. We, data scientists, most often deal with the one-dimensional case, where the “line” is time, and the points are the events of interest — I dare to say.

These are the assumptions of the Poisson process:

  1. The occurrence of one event does not affect the probability of a second event. Think of our seller going on to list another item tomorrow indifferently of having done so already today, or the one from five days ago for that matter. The point here is that there is no memory between events.
  2. The average rate at which events occur, is independent of any occurrence. In other words, no event that happened (or will happen) alters λ, which remains constant throughout the observed timeframe. In our seller example, this means that listing an item today does not increase or decrease the seller’s motivation or likelihood of listing another item tomorrow.
  3. Two events cannot occur at exactly the same instant. If we were to zoom at an infinite granular level on the timescale, no two listings could have been placed simultaneously; always sequentially.

From these assumptions — no memory, constant rate, events happening alone — it follows that 1) any interval’s number of events is Poisson-distributed with parameter λₜ and 2) that disjoint intervals are independent — two key properties of a Poisson process.

A Note on the distribution:
The distribution simply describes probabilities for various numbers of counts in an interval. Strictly speaking, one can use the distribution pragmatically whenever the data is nonnegative, can be unbounded on the right, has mean λ, and reasonably models the data. It would be just convenient if the underlying process is a Poisson one, and actually justifies using the distribution.

The marketplace example: Implications

So, can we justify using the Poisson distribution for our marketplace example? Let’s open up the assumptions of a Poisson process and take the test.

Constant λ

  • Why it may fail: The seller has patterned online activity; holidays; promotions; listings are seasonal goods.
  • Consequence: λ is not constant, leading to overdispersion (mean-to-variance ratio is larger than 1, or to temporal patterns.

Independence and memorylessness

  • Why it may fail: The propensity to list again is higher after a successful listing, or conversely, listing once depletes the stock and intervenes with the propensity of listing again.
  • Consequence: Two events are no longer independent, as the occurrence of one informs the occurrence of the other.

Simultaneous events

  • Why it may fail: Batch-listing, a new feature, was introduced to help the sellers.
  • Consequence: Multiple listings would come online at the same time, clumped together, and they would be counted simultaneously.

Balancing rigour and pragmatism

As Data Scientists on the job, we may feel trapped between rigour and pragmatism. The three steps below should give you a sound foundation to decide on which side to err, when the Poisson distribution falls short:

  1. Pinpoint your goal: is it inference, simulation or prediction, and is it about high-stakes output? List the worst thing that can happen, and the cost of it for the business.
  2. Identify the problem and solution: why does the Poisson distribution not fit, and what can you do about it? list 2-3 solutions, including changing nothing.
  3. Balance gains and costs: Will your workaround improve things, or make it worse? and at what cost: interpretability, new assumptions introduced and resources used. Does it help you in achieving your goal?

That said, here are some counters I use when needed.

When real life deviates from your model

Everything described so far pertains to the standard, or homogenous, Poisson process. But what if reality begs for something different?

In the next section, we’ll cover two extensions of the Poisson distribution when the constant λ assumption does not hold. These are not mutually exclusive, but neither they are the same:

  1. Time-varying λ: a single seller whose listing rate ramps up before holidays and slows down afterward
  2. Mixed Poisson distribution: multiple sellers listing items, each with their own λ can be seen as a mixture of various Poisson processes

Time-varying λ

The first extension allows λ to have its own value for each time t. The PMF then becomes

Where the number of events 𝐾(𝑇) in an interval 𝑇 follows the Poisson distribution with a rate no longer equal to a fixed λ, but one equal to:

More intuitively, integrating over the interval 𝑡 to 𝑡 + 𝑖 gives us a single number: the expected value of events over that interval. The integral will vary by each arbitrary interval, and that’s what makes λ change over time. To understand how that integration works, it was helpful for me to think of it like this: if the interval 𝑡 to 𝑡₁ integrates to 3, and 𝑡₁ to 𝑡₂ integrates to 5, then the interval 𝑡 to 𝑡₂ integrates to 8 = 3 + 5. That’s the two expectations summed up, and now the expectation of the entire interval.

Practical implication 
One may want to modeling the expected value of the Poisson distribution as a function of time. For instance, to model an overall change in trend, or seasonality. In generative model notation:

Time may be a continuous variable, or an arbitrary function of it.

Process-varying λ: Mixed Poisson distribution

But then there’s a gotcha. Remember when I said that λ has a dual role as the mean and variance? That still applies here. Looking at the “relaxed” PMF*, the only thing that changes is that λ can vary freely with time. But it’s still the one and only λ that orchestrates both the expected value and the dispersion of the PMF*. More precisely, 𝔼[𝑋] = Var(𝑋) still holds.

There are various reasons for this constraint not to hold in reality. Model misspecification, event interdependence and unaccounted for heterogeneity could be the issues at hand. I’d like to focus on the latter case, as it justifies the Negative Binomial distribution — one of the topics I promised to open up.

Heterogeneity and overdispersion
Imagine we are not dealing with one seller, but with 10 of them listing at different intensity levels, λᵢ, where 𝑖 = 1, 2, 3, …, 10 sellers. Then, essentially, we have 10 Poisson processes going on. If we unify the processes and estimate the grand λ, we simplify the mixture away. Meaning, we get a correct estimate of all sellers on average, but the resulting grand λ is naive and does not know about the original spread of λᵢ. It still assumes that the variance and mean are equal, as per the axioms of the distribution. This will lead to overdispersion and, in turn, to underestimated errors. Ultimately, it inflates the false positive rate and drives poor decision-making. We need a way to embrace the heterogeneity amongst sellers’ λᵢ.

Negative binomial: Extending the Poisson distribution
Among the few ways one can look at the Negative Binomial distribution, one way is to see it as a compound Poisson process — 10 sellers, sounds familiar yet? That means multiple independent Poisson processes are summed up to a single one. Mathematically, first we draw λ from a Gamma distribution: λ ~ Γ(r, θ), then we draw the count 𝑋 | λ ~ Poisson(λ).

In one image, it is as if we would sample from plenty Poisson distributions, corresponding to each seller.

A negative Binomial distribution arises from many Poisson distributions.
A negative Binomial distribution arises from many Poisson distributions.

The more exposing alias of the Negative binomial distribution is Gamma-Poisson mixture distribution, and now we know why: the dictating λ comes from a continuous mixture. That’s what we needed to explain the heterogeneity amongst sellers.

Let’s simulate this scenario to gain more intuition.

Gamma mixture of lambda.
Gamma mixture of lambda.

First, we draw λᵢ from a Gamma distribution: λᵢ ~ Γ(r, θ). Intuitively, the Gamma distribution tells us about the variety in the intensity — listing rate — amongst the sellers.

On a practical note, one can instill their assumptions about the degree of heterogeneity in this step of the model: how different are sellers? By varying the levels of heterogeneity, one can observe the impact on the final Poisson-like distribution. Doing this type of checks (i.e., posterior predictive check), is common in Bayesian modeling, where the assumptions are set explicitly.

Gamma-Poisson mixture distribution versus homogenous Poisson distribution. Τhe dashed line reflects λ, which is 4 for both distributions.
Gamma-Poisson mixture distribution versus homogenous Poisson distribution. Τhe dashed line reflects λ, which is 4 for both distributions.

In the second step, we plug the obtained λ into the Poisson distribution: 𝑋 | λ ~ Poisson(λ), and obtain a Poisson-like distribution that represents the summed subprocesses. Notably, this unified process has a larger dispersion than expected from a homogeneous Poisson distribution, but it is in line with the Gamma mixture of λ.

Heterogeneous λ and inference

A practical consequence of introducing flexibility into your assumed distribution is that inference becomes more challenging. More parameters (i.e., the Gamma parameters) need to be estimated. Parameters act as flexible explainers of the data, tending to overfit and explain away variance in your variable. The more parameters you have, the better the explanation may seem, but the model also becomes more susceptible to noise in the data. Higher variance reduces the power to identify a difference in means, if one exists, because — well — it gets lost in the variance.

Countering the loss of power

  1. Confirm whether you indeed need to extend the standard Poisson distribution. If not, simplify to the simplest, most fit model. A quick check on overdispersion may suffice for this.
  2. Pin down the estimates of the Gamma mixture distribution parameters using regularising, informative priors (think: Bayes).

During my research process for writing this blog, I learned a great deal about the connective tissue underlying all of this: how the binomial distribution plays a fundamental role in the processes we’ve discussed. And while I’d love to ramble on about this, I’ll save it for another post, perhaps. In the meantime, feel free to share your understanding in the comments section below 👍.

Conclusion

The Poisson distribution is a simple distribution that can be highly suitable for modelling count data. However, when the assumptions do not hold, one can extend the distribution by allowing the rate parameter to vary as a function of time or other factors, or by assuming subprocesses that collectively make up the count data. This added flexibility can address the limitations, but it comes at a cost: increased flexibility in your modelling raises the variance and, consequently, undermines the statistical power of your model.

If your end goal is inference, you may want to think twice and consider exploring simpler models for the data. Alternatively, switch to the Bayesian paradigm and leverage its built-in solution to regularise estimates: informative priors.

I hope this has given you what you came for — a better intuition about the Poisson distribution. I’d love to hear your thoughts about this in the comments!

Unless otherwise noted, all images are by the author.
Originally published at 
https://aalvarezperez.github.io on January 5, 2025.

Shape
Shape
Stay Ahead

Explore More Insights

Stay ahead with more perspectives on cutting-edge power, infrastructure, energy,  bitcoin and AI solutions. Explore these articles to uncover strategies and insights shaping the future of industries.

Shape

AMD agrees to buy World Labs to fill out its AI stack

This is where World Labs fits into AMD’s ecosystem, according to Parv Sharma, Senior Research Analyst at Counterpoint Research. “World Labs builds AI that understands space, where models need to understand geometry, physics and time, unlike LLMs, which understand languages. These world models are used for training in physical AI

Read More »

Why a network digital twin is the missing piece for AI-era operations

The e-book draws an important distinction between two approaches that share the label. One emulates the network by running the actual device firmware against specific test scenarios. The other builds a deterministic mathematical model from the network’s configuration and state, computing all possible forwarding behaviors at once. The guide sums

Read More »

NetScaler admins told to patch critical zero-days in ADC and Gateway now

NetScaler appliances are an important part of many enterprise networks, providing VPN and remote access, load balancing and other application delivery services. Citrix is tracking the two exploited vulnerabilities as CVE-2026-88771 and CVE-2026-88772. It has released fixes in NetScaler ADC and Gateway 14.1-73.37 and later, 13.1-64.23 and later, with corresponding

Read More »

S&P Global: Canadian oil sands output set for record 3.5 million b/d in 2026

Canadian oil sands production is expected to rise for a 25th consecutive year in 2026, reaching a record 3.5 million b/d as operators continue to optimize existing installations, according to S&P Global Energy. The forecast represents an increase of about 100,000 b/d, or 3%, from 2025. S&P Global expects production to reach about 3.9 million b/d by the early 2030s before broadly plateauing under its current outlook. Oil sands production has expanded steadily over the past quarter century. Annual output stood at about 300,000 b/d in 2001 and has increased every year since then except in 2020, when production was affected by the COVID-19 pandemic. Most of the anticipated 2026 growth is set to come from optimization of existing operations rather than major new projects. Much of Canada’s current oil sands capacity was built between 2009-18, while construction of large, new installations has been limited in recent years. S&P Global said, however, that the potential for renewed interest in capacity additions through new construction is creating additional upside to the longer-term outlook. “The Canadian oil sands has proven to be a resilient source of supply despite periods of low oil prices, regional price volatility and uncertainty over future Canadian energy and climate policy,” said Kevin Birn, chief Canadian oil markets analyst at S&P Global Energy. “The question today is not whether the oil sands will continue to grow, but rather how much additional growth could come should new projects once again come forward.” Factors contributing to that outlook include announced plans for expanded pipeline export capacity, greater clarity, reduction, and extension of carbon pricing through 2040, commitments to accelerate reviews of projects considered to be in the national interest, and potential changes to fiscal terms for new oil sands projects. S&P Global also said Canadian energy production is increasingly being

Read More »

Insights: Prioritizing process safety management across the refining industry (Pt. II)

In this Insights episode of the Oil & Gas Journal ReEnterprised podcast, downstream editor and lead reporter Robert Brelsford concludes the conversation on process safety management with Joe Barnes, principal of Barnes’ Engineering Consultants and a veteran oil and gas operations, maintenance, reliability, and major-projects leader with more than 30 years of industry experience. Barnes outlines five priorities for senior leaders and plant managers: maintaining strong operating procedures and management-of-change practices; routinely auditing lockout/tagout procedures; providing effective operator and supervisor training; and preventing deferred maintenance, particularly on critical equipment. He also emphasizes the importance of leadership visibility in the field, encouraging executives and plant managers to regularly engage with operating personnel and understand the condition of equipment and processes firsthand. During the conversation, Robert and Joe examine the tension between profitability, market pressures, and the investments required to operate safely. Barnes explains why strong financial performance should support investment in equipment, maintenance, training, and people rather than encourage short-term decisions that could increase risk. He also discusses how management should evaluate efficiency measures, staffing changes, maintenance deferrals, turnaround modifications, and capital reductions while keeping process and personal safety central to decision-making. The two consider how facilities can demonstrate that process safety systems are working in practice through dedicated internal audits, periodic third-party reviews, corrective-action follow-through, and the use of real-world incident scenarios in operator training. Barnes explains how investigation findings can be converted into case-based exercises that help workers recognize hazards, assess risks, and develop the right responses before returning to the field. He also recommends reviewing procedures, emergency-response plans, and process hazards periodically to help ensure they remain current and effective. Robert and Joe further explore how leaders can create an environment in which employees and contractors feel comfortable reporting weak signals, near misses, and other bad news. Barnes

Read More »

Kimmeridge: US shale oil reserve replacement weakens as gas remains abundant

US shale oil producers are finding it increasingly difficult to replace reserves even as operating efficiency improves, while natural gas resources remain comparatively abundant, Kimmeridge, an alternative asset manager focused on the energy sector, said in a new report. In the report titled “Shale’s Golden Years, Part II: The Cost of Aging,” Kimmeridge said cumulative oil reserve replacement since 2019 has averaged about 95%, compared with 126% for natural gas. Well-level data show a similar split, with oil recovery per foot generally declining over the past decade while gas productivity has remained broadly flat to improving. The deterioration comes despite majort cost and efficiency gains. Since 2018, SG&A expense per barrel of oil equivalent (boe) has fallen about 48%, interest expense 46%, and exploration expense 71%. Operators also have drilled longer laterals and increased drilling speeds. Even so, Kimmeridge’s 3-year, value-weighted recycle ratio for the US E&P sector fell to 167% in 2025 from 184% in 2019, despite higher revenue per boe. Oil-weighted producers generated a 164% recycle ratio in 2025 versus 186% in 2019, while gas-weighted producers improved to 179% from 165%. Reserve additions at oil-focused companies also are becoming gassier. Oil represented 41% of their reserve additions in 2025, compared with about 50% of current production. Kimmeridge said the conventional 6 Mcf-to-1 boe conversion can obscure that shift by giving lower-value gas the same energy-equivalent replacement credit as oil. Kimmeridge said weaker oil reserve replacement could reduce the responsiveness of US shale supply to higher prices over time, providing structural support for WTI and strengthening the case for renewed oil exploration. Natural gas faces the opposite backdrop. Efficient gas reserve replacement and rising associated-gas output point to continued supply abundance, potentially weighing on Henry Hub and increasing the value of LNG-linked sales, transportation, and other downstream exposure.

Read More »

TotalEnergies takes FID for Absheron full field development offshore Azerbaijan

TotalEnergies has taken final investment decision (FID) for the full field development of Absheron gas and condensate field in the Caspian Sea offshore Azerbaijan. Absheron field, which lies 100 km southeast of Baku, is estimated to hold about 140 bcm of recoverable gas reserves, TotalEnergies said in a release Sept. 26. The first development phase, with a production capacity of 1.5 bcmy of gas and 12,000 b/d of condensate, came on stream in 2023. Absheron Full Field development is expected to increase overall field production to 6 bcmy and 47,000 b/d of condensate. Startup is expected in 2029. Gas will be supplied to the domestic market and exported to Türkiye through the existing gas infrastructure connecting Azerbaijan to the European gas market. Gas will be produced by 4 subsea wells, transported to shore through a subsea pipeline equipped with advanced automation solutions and processed in a new onshore plant, fully electrified and designed to minimize energy consumption and greenhouse gas emissions, TotalEnergies said. The scope 1 & 2 greenhouse gas emissions intensity of the project is below 4 kg CO2e/boe. TotalEnergies is operator of the project with 35% interest. Partners are State Oil Co. of the Republic of Azerbaijan (SOCAR, 35%) and XRG, a unit of ADNOC (30%).  

Read More »

Trump Administration Mitigates Blackout Risks by Keeping Colorado Coal Plant Online

WASHINGTON—U.S. Secretary of Energy Chris Wright today issued an emergency order to keep a Colorado coal plant operational to ensure Americans maintain access to affordable, reliable, and secure electricity. The order directs Tri-State Generation and Transmission Association (Tri-State), Southwest Power Pool (SPP), Platte River Power Authority, Salt River Project, PacifiCorp, and Public Service Company of Colorado to take all measures necessary to ensure that Unit 1 at the Craig Station in Craig, Colorado is available to operate. For the duration of this Order, SPP is directed to take every step to employ economic dispatch of Craig Unit 1 to mitigate the risk of blackouts and minimize costs to taxpayers. Unit 1 of the coal plant was originally scheduled to shut down at the end of 2025, but in December 2025 and twice in 2026, Secretary Wright issued emergency orders directing Tri-State and the co-owners to ensure that Craig Unit 1 remains available to operate. “For over the past decade, state and federal leaders have harmed Coloradans’ wallets and energy security with efforts to force reliable generation off the grid,” said Secretary Wright. “The Trump Administration will continue taking action to ensure we don’t lose critical generation sources. Americans deserve access to affordable, reliable, and secure energy to power their homes all the time, regardless of whether the wind is blowing or the sun is shining.” Thanks to President Trump’s leadership, coal generating plants across the country are being saved from premature retirement. For example, since 2025 more than 17 gigawatts of coal power electricity generation were saved from going offline. The availability of Craig Unit 1 to operate will continue to be an asset to maintain reliability in the Western Electricity Coordinating Council (WECC) Rocky Mountain region and is necessary to address elevated reliability risks in the region during atypical weather and reduce

Read More »

MOL wraps repairs on major unit at Hungarian refinery

Hungary’s MOL Group has completed repairs to a main crude processing unit that suffered major damage resulting from a late-2025 fire at its 8.1-million tonnes/year (163,000-b/cd) refinery along the Danube River in Százhalombatta, near Budapest. Repair-related construction works amounting to 18 billion forints (Hun.; US$57.1 million) were completed as of Sept. 22 on the refinery’s atmospheric-vacuum distillation (AV-3) unit, paving the way for the unit’s gradual restart once construction equipment is removed, MOL said in a release. Completed according to schedule, the year-long comprehensive repair project entailed the demolition and rebuilding of the unit’s reinforced concrete structure, as well as restoration works on all unit equipment impacted by the fire, including related control and electrical systems, the operator said. Alongside mechanical repairs, the project also included: Installation of 27 new, high-capacity pumps. Replacement of 8 km of piping. Installation of 100 km of electrical cables. With construction activities now wrapped, MOL said the refinery is preparing for the unit’s gradual restart following a series of next steps that will include: Flushing of unidentified unit systems. Executing pressure tests. Performing functional tests of the process control and safety equipment. Connecting of all associated auxiliary power supplies. Coordinating operation of the repaired AV-3 unit with associated plants of the refinery. Once all preliminary restart activities and complementary operational safety inspections and tests have been completed, crude throughputs will be reintroduced into the unit for phased restart of production activities beginning in October, MOL said. Requisite repairs to AV-3 follow a fire that broke out in the unit on Oct. 20, 2025, which led to a temporary shutdown of the refinery and a 50% drop in crude distillation capacity at the site following the incident, according to the operator’s 2025 annual report to investors. While plants not affected by the fire were quickly

Read More »

Q3 Executive Roundtable Recap

For Data Center Frontier’s Q3 2026 Executive Roundtable, three industry leaders examined a question increasingly central to the AI infrastructure buildout: What happens when data centers are asked to become larger, denser and faster at the same time? Across three discussions, a consistent theme emerged. AI is not simply increasing the amount of infrastructure required to support the modern data center. It is exposing assumptions that were easier to tolerate at lower densities, expanding the boundaries of what operators must consider mission-critical, and making the interaction between systems increasingly important to overall resilience. That begins with density. As the value and power concentrated in individual racks rises, traditional approaches to redundancy, monitoring and risk mitigation can leave less room for error. Resilience can no longer be measured simply by installed capacity or the presence of backup equipment. Operators increasingly need to understand how electrical, thermal and control systems behave together under dynamic AI workloads — and how quickly the facility can respond and recover when something goes wrong. The same shift is broadening the definition of critical infrastructure. Power generation, UPS systems and network connectivity remain fundamental, but energy storage, liquid cooling, leak detection, controls, monitoring and the interfaces connecting them are becoming part of the same reliability equation. A component can perform exactly as designed while the larger system still fails if coordination, communications or control logic break down. And all of this is happening while the market is demanding faster deployment. Standardization, modular construction, factory integration and earlier modeling can legitimately compress project schedules. But the Q3 panelists drew a clear distinction between eliminating unnecessary time and eliminating rigor. As infrastructure becomes more tightly coupled, commissioning, integrated systems testing, operational visibility and system-level validation may need to become more thorough precisely because projects are moving faster. Taken together,

Read More »

DCF Tours: Inside CoolIT, Where AI Liquid Cooling Goes to Scale

How Long Can Single-Phase Go? The Liquid Lab also makes clear that CoolIT is not treating today’s architecture as permanent. The company’s R&D operation includes CNC machining, 3D printing, skiving equipment and friction stir welding, allowing engineers to move quickly from CAD designs to physical prototypes. Some work is aimed several processor generations ahead. CoolIT is also experimenting with two-phase thermal technologies. That does not mean the company expects two-phase cooling to displace single-phase DLC wholesale. Robison sees the technologies as potentially complementary. Two-phase techniques can be particularly effective for moving heat from localized areas, using approaches such as vapor chambers and heat pipes. But operating an entire data center cooling loop through repeated phase changes introduces another set of system-engineering challenges. CoolIT’s position is that single-phase cooling still has significant room to advance through better geometries, flow management and system design, while two-phase technologies may emerge where they provide a specific thermal advantage. That is a more useful way to think about the cooling transition than searching for a single architecture that wins outright. AI servers are becoming collections of thermal problems rather than a single thermal problem. Processors, memory, networking and storage may ultimately require different cooling approaches even within the same system. The thermal architecture is likely to become more diverse as density rises. Cooling Becomes Infrastructure Walking through the CoolIT campus in Calgary, the most striking feature was not any single cold plate, manifold or CDU. It was the amount of infrastructure now required to develop and validate the cooling infrastructure itself. A cold plate begins as a carefully engineered flow path measured in millimeters. Several steps later, that component has become part of a megawatt-scale thermal system involving pumps, controls, manifolds, piping, facility water and field technicians. And before that system reaches a data center,

Read More »

Rewiring global capability centers for the AI era

When a global capability center (GCC) underdelivers, the diagnosis is usually people: wrong hires, wrong scope, not enough seniority. It’s rarely the honest answer. More often the center was wired like a branch office and asked to behave like a headquarters. The GCC has evolved from an offshore cost play to a strategic extension of HQ, owning engineering, product, and, increasingly, the AI build. This is no longer solely for the Fortune 500. Leaner centers of 50–200 people, as well as ‘GCC-as-a-service’ and managed models, put it within reach of many US midmarket companies. Demand for AI is accelerating this trend further. India alone now has more than 2,000 GCCs, generating $98.4 billion in revenue in the fiscal year 2026. “A GCC is never about cost effectiveness, it’s about tapping the best talent to take enterprises to the next technological orbit. The GCC model is shifting to intellectual arbitrage,” says Murali Krishnan, AVP & Head of Business – Enterprise Network at Tata Communications. With lower barriers to entry, midmarket companies are looking to tap into this opportunity. However, this size of business tends to carry a domestic, branch-office playbook into a GCC and wire it accordingly. While that network was good enough for a branch office, it actively limits what a GCC can do and can limit their return on investment. Midmarket companies looking to tap into this opportunity face a structural mismatch. Existing networks connect offices to a data center inside one country, with bandwidth sized accordingly. A GCC introduces AI workloads across several clouds and two continents, and the branch office network reaches its design limits. What AI-driven workflows demand from the network The growth in GCCs, alongside the rapid adoption of AI, has increased strain on the network, driving demand for high-performance connectivity across geographies. Model training, data-pipeline engineering,

Read More »

Google’s first prototype satellite is going up, kicking off its space-based data center project

Google announced its Project Suncatcher space-based data center plan last November. The goal is to take advantage of the unlimited sunlight of space and to minimize the impact of data centers here on Earth. Next week, the first satellite is going up on the SpaceX Transporter-18 mission, Google announced yesterday. The prototype satellite will test how Google’s AI chips—its Tensor Processing Units—perform in space, says Travis Beals, Google’s Senior Director, Paradigms of Intelligence, in the announcement. Google has already conducted some testing here on Earth. The TPU chips were able to handle the level of vibration and acceleration that they would see during the launch, and be able to survive a bigger radiation dose than they would receive during a five-year space mission. In addition, the team has tested a cooling system—a combination of heat pipes and radiators—in a thermal vacuum chamber that simulates space.

Read More »

Anthropic, OpenAI Keep Expanding the AI Data Center Map — and the Financing Gets Harder

September has offered one of the clearest pictures yet of what the frontier AI race looks like when translated from models and tokens into physical infrastructure. Anthropic has moved aggressively to lock down dedicated compute in the United States while establishing its first major data center foothold in Australia. OpenAI, meanwhile, has expanded into Malaysia through Nvidia-backed Firmus as the financing behind its much larger infrastructure ambitions continues to grow more complicated. All in all, the developments suggest that competition between the leading AI labs is entering another phase. Securing GPUs remains essential, but the harder problem is increasingly assembling the entire chain around them: land, power, cooling, networks, project finance and counterparties capable of delivering capacity measured in hundreds of megawatts — and increasingly gigawatts. That distinction is important for the data center industry. The AI infrastructure story is no longer simply about projected demand. It is increasingly about which commitments can actually become operating megawatts. Anthropic’s $45 Billion Bet Gets More Concrete The most revealing new detail came not from Anthropic itself, but from Nscale. The Nvidia-backed AI infrastructure provider filed for a U.S. initial public offering on Sept. 18, providing new financial and technical detail around a massive compute agreement first reported in August. Nscale’s SEC filing says it entered four GPU services agreements with Anthropic on Aug. 25 that could generate approximately $44.6 billion in aggregate payments. The agreements call for Nscale to provide Anthropic with dedicated infrastructure built around Nvidia Vera Rubin NVL72 systems at the company’s planned Monarch Compute Campus in Mason County, West Virginia. The deployments are structured in four tranches with multiyear service terms. Reuters previously reported the agreement at roughly $45 billion over six years, covering about 460 MW of compute capacity at Monarch. (The agreement follows Anthropic’s $19 billion, 401-MW

Read More »

Amazon-Generac Deal Puts Backup Power in the AI Infrastructure Spotlight

Amazon has struck a long-term supply agreement with Generac for backup generators supporting its data center buildout, tying one of the cloud industry’s largest infrastructure programs to a manufacturer that has been rapidly expanding into the hyperscale power market. Under the agreement disclosed in a Sept. 16 regulatory filing, Generac expects initial deliveries to Amazon totaling approximately $2.4 billion during 2027 and 2028. The commercial relationship could ultimately involve as much as $8 billion in qualifying generator purchases. The agreement also gives Amazon an equity interest in Generac’s success. Generac issued Amazon.com NV Investment Holdings a warrant to acquire as many as 1.69 million Generac shares at an exercise price of approximately $200.93 per share. About 308,000 shares vested when the agreement was signed, with additional tranches vesting as Amazon’s purchases increase. The warrant remains exercisable through September 2033. The distinction is important: the frequently cited $8 billion figure represents potential cumulative payments by Amazon for backup power generators, rather than an $8 billion equity investment. The maximum warrant covers roughly $340 million of Generac stock at the stated exercise price. CNBC first highlighted the equity component of the transaction, reporting that Generac shares surged more than 40% in extended trading following disclosure of the agreement. The shares ultimately gained about 18% during the following regular trading session. Generac Was Already Scaling for the Data Center Market For the data center industry, however, the more consequential part of the transaction may be the size and duration of Amazon’s equipment commitment. Generac has spent much of the past two years positioning itself as an alternative large-megawatt generator supplier as AI infrastructure development puts pressure on established power-equipment supply chains. DCF previously examined Generac’s push into hyperscale backup power, including its effort to shorten generator lead times and support campuses requiring hundreds

Read More »

Microsoft will invest $80B in AI data centers in fiscal 2025

And Microsoft isn’t the only one that is ramping up its investments into AI-enabled data centers. Rival cloud service providers are all investing in either upgrading or opening new data centers to capture a larger chunk of business from developers and users of large language models (LLMs).  In a report published in October 2024, Bloomberg Intelligence estimated that demand for generative AI would push Microsoft, AWS, Google, Oracle, Meta, and Apple would between them devote $200 billion to capex in 2025, up from $110 billion in 2023. Microsoft is one of the biggest spenders, followed closely by Google and AWS, Bloomberg Intelligence said. Its estimate of Microsoft’s capital spending on AI, at $62.4 billion for calendar 2025, is lower than Smith’s claim that the company will invest $80 billion in the fiscal year to June 30, 2025. Both figures, though, are way higher than Microsoft’s 2020 capital expenditure of “just” $17.6 billion. The majority of the increased spending is tied to cloud services and the expansion of AI infrastructure needed to provide compute capacity for OpenAI workloads. Separately, last October Amazon CEO Andy Jassy said his company planned total capex spend of $75 billion in 2024 and even more in 2025, with much of it going to AWS, its cloud computing division.

Read More »

John Deere unveils more autonomous farm machines to address skill labor shortage

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Self-driving tractors might be the path to self-driving cars. John Deere has revealed a new line of autonomous machines and tech across agriculture, construction and commercial landscaping. The Moline, Illinois-based John Deere has been in business for 187 years, yet it’s been a regular as a non-tech company showing off technology at the big tech trade show in Las Vegas and is back at CES 2025 with more autonomous tractors and other vehicles. This is not something we usually cover, but John Deere has a lot of data that is interesting in the big picture of tech. The message from the company is that there aren’t enough skilled farm laborers to do the work that its customers need. It’s been a challenge for most of the last two decades, said Jahmy Hindman, CTO at John Deere, in a briefing. Much of the tech will come this fall and after that. He noted that the average farmer in the U.S. is over 58 and works 12 to 18 hours a day to grow food for us. And he said the American Farm Bureau Federation estimates there are roughly 2.4 million farm jobs that need to be filled annually; and the agricultural work force continues to shrink. (This is my hint to the anti-immigration crowd). John Deere’s autonomous 9RX Tractor. Farmers can oversee it using an app. While each of these industries experiences their own set of challenges, a commonality across all is skilled labor availability. In construction, about 80% percent of contractors struggle to find skilled labor. And in commercial landscaping, 86% of landscaping business owners can’t find labor to fill open positions, he said. “They have to figure out how to do

Read More »

2025 playbook for enterprise AI success, from agents to evals

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More 2025 is poised to be a pivotal year for enterprise AI. The past year has seen rapid innovation, and this year will see the same. This has made it more critical than ever to revisit your AI strategy to stay competitive and create value for your customers. From scaling AI agents to optimizing costs, here are the five critical areas enterprises should prioritize for their AI strategy this year. 1. Agents: the next generation of automation AI agents are no longer theoretical. In 2025, they’re indispensable tools for enterprises looking to streamline operations and enhance customer interactions. Unlike traditional software, agents powered by large language models (LLMs) can make nuanced decisions, navigate complex multi-step tasks, and integrate seamlessly with tools and APIs. At the start of 2024, agents were not ready for prime time, making frustrating mistakes like hallucinating URLs. They started getting better as frontier large language models themselves improved. “Let me put it this way,” said Sam Witteveen, cofounder of Red Dragon, a company that develops agents for companies, and that recently reviewed the 48 agents it built last year. “Interestingly, the ones that we built at the start of the year, a lot of those worked way better at the end of the year just because the models got better.” Witteveen shared this in the video podcast we filmed to discuss these five big trends in detail. Models are getting better and hallucinating less, and they’re also being trained to do agentic tasks. Another feature that the model providers are researching is a way to use the LLM as a judge, and as models get cheaper (something we’ll cover below), companies can use three or more models to

Read More »

OpenAI’s red teaming innovations define new essentials for security leaders in the AI era

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More OpenAI has taken a more aggressive approach to red teaming than its AI competitors, demonstrating its security teams’ advanced capabilities in two areas: multi-step reinforcement and external red teaming. OpenAI recently released two papers that set a new competitive standard for improving the quality, reliability and safety of AI models in these two techniques and more. The first paper, “OpenAI’s Approach to External Red Teaming for AI Models and Systems,” reports that specialized teams outside the company have proven effective in uncovering vulnerabilities that might otherwise have made it into a released model because in-house testing techniques may have missed them. In the second paper, “Diverse and Effective Red Teaming with Auto-Generated Rewards and Multi-Step Reinforcement Learning,” OpenAI introduces an automated framework that relies on iterative reinforcement learning to generate a broad spectrum of novel, wide-ranging attacks. Going all-in on red teaming pays practical, competitive dividends It’s encouraging to see competitive intensity in red teaming growing among AI companies. When Anthropic released its AI red team guidelines in June of last year, it joined AI providers including Google, Microsoft, Nvidia, OpenAI, and even the U.S.’s National Institute of Standards and Technology (NIST), which all had released red teaming frameworks. Investing heavily in red teaming yields tangible benefits for security leaders in any organization. OpenAI’s paper on external red teaming provides a detailed analysis of how the company strives to create specialized external teams that include cybersecurity and subject matter experts. The goal is to see if knowledgeable external teams can defeat models’ security perimeters and find gaps in their security, biases and controls that prompt-based testing couldn’t find. What makes OpenAI’s recent papers noteworthy is how well they define using human-in-the-middle

Read More »