Greene King
"I’ve been wrong more than I’ve been right”: Greene King’s John Rowley on proving the case for experimentation
Greene King, the pub company and brewer, already ran on data. Two years ago, it started testing whether the data was telling it the truth.
The premise was straightforward: rather than use data simply to inform a decision, test whether the decision stands up. The harder part was proving the commercial case, including why finding out what not to do can be almost as valuable as finding the next source of growth.
Q. You were already using data extensively at Greene King. What was missing?
Being data-driven has its own flaws. You find a data point and cognitive bias can do the rest. What we weren’t doing was scientifically stress-testing what we thought we knew. We wanted to ask: is that still true? To what degree? How far can we push the knowledge? And how much does it matter compared with everything else we think we know? Experimentation gave us a way to answer those questions.
Q. What did you have to prove to get the board behind it?
We had buy-in for the concept and permission to run a proof of concept. We judged it on four things: commercial value, customer benefit, what the company could learn and what it meant for colleagues.
The commercial case had to be concrete, so we translated bookings into pounds and pence at revenue and margin level and built that into the cost-benefit model. But we also wanted to know whether experimentation could improve the customer journey, give us insight we didn’t already have and create opportunities for colleagues to learn and develop.
From concept to completed proof of concept took around eight to 12 months. By the end, we had a positive cost-benefit case to take back to the board.
Q. Was there a moment when the case was made?
There was an interesting tension. Our early tests produced lots of small, incremental gains. But what everyone naturally wanted to see was one big number. That isn’t generally how CRO works. I call it ‘death by a thousand tests’: lots of relatively small improvements that accumulate into something significant versus “big bang” testing in which you hope for a single test to provide significant uplift.
So, we partnered with our booking’s product team and split-tested one of its major launches. The impact was eight to 10 times larger than anything we’d seen from a single test. When we extrapolated the result across a full year and all our venues, it delivered more than 100% of the commercial benefit we needed to prove. That was the moment we knew we had the case.
Q. What has the programme delivered?
In the time I’ve been with the organisation, we’ve had significant booking growth year-over-year, , with experimentation a key contributor alongside other digital initiatives.
The other half of the story is what we didn’t launch. Testing has stopped us putting changes live that could have cost us revenue. We put a commercial value against that avoided loss, and it’s almost as significant as the incremental revenue generated through bookings. It’s harder to demonstrate because it never appears on the P&L, but commercially it matters.
Q. Does that require people to become more comfortable with being wrong?
Yes. We try not to describe tests as winners or failures. If a test disproves an idea, that’s still a learning. I’m terrible at this myself, I still ask whether something worked, but the language matters. People need to be able to put an idea forward and be proved wrong without that being seen as a failure.
And knowing what happened isn’t enough. We want to understand why, and whether the result changes by channel, customer segment, brand, proposition, timing or stage in the journey. A test can tell you one thing in one context and something quite different in another. Getting to the real learning can take a lot of analysis.
Q. Knowing what you know now, how would you build the case differently?
I’d hypothesise the big bets before the thousand small ones. When you’re trying to secure investment, find the larger opportunity that can demonstrate the value. Once you’ve proved the case, you can settle into the steady rhythm of smaller, incremental tests.
From there, start simple because the learning compounds. Take photography: do people, products or places drive more bookings? If people perform better, is it one person, a couple or a group? If groups perform better, are they watching sport, drinking or eating?
Then you start combining those findings with copy, calls to action and placement, across different journeys and customer segments. A simple question becomes a large, complex data set very quickly. You need to think about testing in a four-dimensional way, which is hard and uncomfortable to begin with.
Q. Can AI change the pace of experimentation?
AI in experimentation isn’t entirely new. Multi-armed bandit tests already use machine learning to adjust the proportion of traffic going to different variants, balancing statistical significance with trading performance. What’s changed is the popularity of AI and the wider understanding of what it can do.
Generative AI opens something different. It can help us create and deploy experiments faster, produce more creative variations to test and get more insight from the results. Ultimately, it could also help us combine what we learn through experimentation with individual customer preferences.
We’re extending our experimentation capabilities and believe AI can help us manage greater volume of tests, as well as increase the speed of deep analysis.
The interesting part is how you combine AI with human thinking. AI is very good at analysing complex, multidimensional data; humans are still better at coming up with the creative hypotheses worth testing. Used together, AI is already increasing the pace of experimentation, and I think that will only accelerate.