Sustainability Simplified (publisher of CSRD Simplified)

Sustainability Simplified (publisher of CSRD Simplified)

Profitable Sustainability

[PS4] How to test sustainability initiatives before scaling them

Profitable Sustainability, article 4: Testing sustainability ideas before betting the balance sheet

Lars Wullink's avatar
Lars Wullink
Oct 11, 2026
∙ Paid

Controlled test campaigns for profitable sustainability

The first three articles in this series established how to run sustainable operations: designing the compounding flywheel in [PS1], protecting new ideas in [PS2], and hitting quarterly milestones in [PS3].

But having the right goals does not prevent bad capital decisions.

Many sustainability investments burn capital on general faith rather than verified return.

Watch how capital gets allocated in almost any enterprise. If a plant manager wants €250,000 for a new packaging conveyor, they must produce a five-year DCF model, vendor quotes, and verified downtime logs. Every euro must justify itself.

But when an enterprise commits to a sustainability initiative, decision-making often flips from operational rigor to speculative assumptions.

Leadership approves a corporate-wide eco-packaging overhaul or a green product line based on high-level sustainability targets and assumed green premiums. Yet research from Bain & Company shows that only 2 percent of corporate sustainability initiatives actually deliver their intended business results.¹ Companies roll out programs across their commercial footprint without running a control test. And then they discover that customer willingness-to-pay is flat and that the projected savings don’t materialize.

The initiative does not fail because the engineering was flawed. It fails because the company scaled an unvalidated hypothesis across its entire balance sheet.

This is the failure mode that advertising pioneer Claude Hopkins diagnosed in 1923 in Scientific Advertising.²

Before Hopkins published his principles, agencies treated advertising as an unmeasured art form. They spent fortunes on general publicity and clever slogans, while assuming that visibility would automatically convert into profit.

Hopkins called this reckless gambling.

He introduced a discipline anchored in small-scale testing, tracked coupons, and strict unit economics. His core insight holds value for sustainability leaders: you do not need to guess what works, and you should never gamble shareholder capital on unproven assumptions. You can isolate and test the core economic assumptions of many operational and commercial initiatives on a bounded scale before committing balance-sheet capital.

In this article, you’ll learn:

✅ Hopkins’ four rules: How Claude Hopkins’ testing principles translate into operational filters

✅ The four-step testing engine: How to validate unit economics and control baselines before CapEx allocation

✅ The transfer gate to the flywheel: How controlled testing bridges early exploration to compounding savings

✅ The Opower case study: How randomized controlled trials proved behavioral efficiency across 600,000 households

By the end, you’ll have a practical testing method to validate sustainability initiatives before committing balance-sheet capital.


What Scientific Advertising means for sustainability

Hopkins built his reputation by turning marketing from an intuitive craft into a measurable science. David Ogilvy later wrote that nobody should be allowed to touch advertising until they had read Hopkins seven times.⁴ When we strip away the 1920s print media context, the book establishes four foundational rules that apply directly to corporate sustainability:

  1. Aim for direct return, not general publicity: The sole purpose of an expenditure is to produce a measurable result.³ Reject the notion that spend should be justified by vague brand elevation. For sustainability teams, this means retiring the excuse that an initiative pays off in unmeasurable goodwill. Every project must prove its return on the P&L, either as lower operating expenses or traceable new revenue.

  2. Treat every initiative like a salesman on commission: Evaluate proposals by asking a simple question: would a personal salesman talk this way to a buyer?³ A good salesman does not recite poetry; they state facts, address customer pain points, and quote prices. Sustainable products succeed when they solve operational headaches, lower maintenance costs, or improve durability, not when they broadcast generic green intentions.

  3. Run test campaigns before national spend: Never roll out an expensive corporate campaign based on boardroom debates. Hopkins placed test advertisements in two or three selected towns, tracked the response, and scaled only when the unit economics were proven.⁶ In corporate operations, facility managers should run isolated, controlled micro-tests before asking for enterprise-wide capital.

  4. Be specific because generalities build zero trust: General claims roll off human understanding like water off a duck’s back.⁷ Saying a product is the best or most economical persuades nobody. Stating that a process cuts electricity draw by 14 percent or reduces material waste by 82 grams commands immediate attention.


The testing engine: four steps to validate sustainability spend

To bring scientific validation into your sustainability workflow, apply these four steps:

Step 1 – Treat every initiative as a salesman on commission

In many companies, sustainability proposals bypass the financial discipline applied to core operations. Research by the NYU Stern Center for Sustainable Business published in Harvard Business Review found that because corporate finance and sustainability teams operate in silos, accounting systems rarely track the operational return on sustainability investments.10 Instead, projects are frequently justified through broad corporate responsibility narratives and brand optics.

When spend is approved on general goodwill rather than tracked return, it inevitably triggers executive skepticism during the next cost-cutting cycle.

Hopkins anchored his core selling rule in human self-interest: never expect a buyer to act out of charity, because people buy only to serve themselves.³ The same principle governs internal corporate proposals. Plant directors and procurement leads will not adopt an eco-design to do the sustainability team a moral favor. They adopt it when the initiative solves an immediate operational friction.

In [PS2], we looked at how to protect new ideas while testing if they can work. But when an idea asks for money from the core operating budget, the rules change. It must now prove that it pays for itself.

When you treat an initiative like a commissioned salesperson, you demand that it justify its existence through direct operational contribution.³ Does the new variable-frequency drive reduce pump electricity consumption enough to service its capital cost? Does lightweight secondary packaging cut freight weight sufficiently to offset retooling expenses?

If an initiative cannot articulate its direct economic return (through lower operating costs, avoided fees, or traceable revenue), it does not belong in the core operating budget. It must either remain in development to improve its economics, or be shelved.

Practical focus: Require a written investment case before allocating test capital. The proposal must state its exact return mechanism on a single page, identifying whether cash flows come from reduced material spend, lower utility consumption, avoided waste tipping fees, or verified price premiums. While sustainable offerings can unlock commercial growth over time [PS1], you must never assume customer willingness-to-pay on faith. Treat customer demand as an unverified hypothesis until tracked purchasing tests prove that buyers will pay the required premium.

Step 2 – Run bounded test campaigns with a contemporaneous control

Corporate sustainability proposals typically get stuck in debate. Engineering and finance teams spend months arguing around meeting room tables over vendor spreadsheets and theoretical payback curves. The CapEx committee either shelves the proposal because the risk feels unproven, or approves an expensive multi-site rollout based on paper assumptions.

The alternative is Hopkins’ operational test campaign: stop arguing and let the plant deliver the data.⁶ Rather than installing heat-recovery systems across twenty manufacturing plants, install the equipment on a single production line, a single distribution route, or a small cluster of retail stores. In the language of [PS3], manage this ninety-day validation trial as an aspirational OKR: a bounded experiment where real performance is observed without risking the core operating budget.

Yet an operational test works only if you isolate cause and effect. In mail-order advertising, every print ad carried a keyed coupon to trace incoming purchases back to their exact source.⁵ In corporate sustainability, the equivalent of a keyed coupon is a contemporaneous control group: testing a change in one area while keeping an identical one untouched at the exact same time.

Many green initiatives declare victory simply because plant energy consumption fell during the trial period. Yet that drop may have been caused by mild outdoor weather, lower shift volumes, or changes in product mix. Without a control baseline established before you launch, you have no way to prove that the savings came from your investment.

To isolate cause and effect, run treatment and control groups side by side:

  • Fleet logistics: Fit ten delivery trucks with aerodynamic fairings while leaving ten identical trucks on the same route unaltered.

  • Process engineering: Apply a new chemical bath to two plating lines while running two adjacent lines on the legacy formulation.

  • Facility operations: Adjust thermostat setpoints in three regional offices while leaving three demographically similar offices on standard settings.

Comparing the two groups proves the real difference the change made. If the test fails, your loss is bounded. If it succeeds, you hold verified operational data that gives capital committees the financial proof they require.

Practical focus: Cap test budgets at what a local site manager can approve on their own signature. Never launch a validation test without establishing a contemporaneous control group. If you cannot isolate a physical control group, establish a weather-normalized, production-weighted historical baseline to benchmark performance.

Step 3 – Measure specific cost per result

Generic claims fail because listeners recognize them as unsubstantiated praise.⁷ Words like "green", "circular", and "sustainable" are dismissed as empty slogans. They fail to persuade internal capital committees, and they fail to convince procurement teams.

Regulatory enforcement is also closing the door on vague claims. In the European Union, Directive (EU) 2024/825 on empowering consumers for the green transition bans generic environmental claims unless companies demonstrate recognized excellent environmental performance.⁸ From 27 September 2026, statements such as “environmentally friendly” or “climate neutral” that lack verified, specific evidence will carry severe legal liability.⁸

Specificity wins on both fronts. Hopkins showed that profitable testing requires calculating the exact cost per result.⁵ In sustainability, that means calculating the cost per kilowatt-hour saved, the cost per metric ton avoided, or the cost per kilogram of diverted scrap. Instead of claiming a product is “responsibly manufactured”, state that the factory uses 42 percent less process water per unit than the regional industry average. Specificity conveys competence, satisfies legal scrutiny, and proves the unit economics of the investment.

Practical focus: Audit project presentations and external marketing claims. Remove all unsupported adjectives. Replace every general claim with a specific physical metric, the measurement baseline, and the reporting period.

Step 4 – Apply the scale-or-stop gate

A scientific testing engine only works if leadership enforces an objective decision gate. Without a predefined threshold, tests drift into what Hopkins called the danger of endless experimentation, becoming an organizational parking lot for unresolved ideas.

Before launching any validation test, establish the proof-of-value threshold:

  • Scale: If the test hits its target cost per result and clears its payback hurdle, advance the initiative through the capital gate into core operations as a committed OKR.

  • Stop: If the test fails to deliver the expected unit savings after ninety days, shut it down. Bounding your losses protects the balance sheet and frees capital for the next test.

This gate also protects companies from Hopkins’ warning: “Let us know the cost of our pride.”³ Corporate sustainability frequently falls into the vanity trap: high-visibility showcase projects, such as shaded rooftop solar arrays on urban headquarters or lobby art made from ocean plastic. These projects generate attractive photos for annual reports, but their return on capital is negligible.

Showcase projects are not forbidden, but they must clear an honest accounting test. If leadership chooses to build a visual display for brand visibility, record it as a corporate communications expense. Never allow vanity projects to masquerade as operational investments, where their weak financial return drags down the credible initiatives that feed your economic engine.

Practical focus: Fix the scale-or-stop threshold before the test launches. Scale what proves its payback into core operations, terminate what fails, and classify executive pride displays strictly as brand marketing.


How this feeds the sustainability flywheel

A disciplined testing engine connects the innovation engine from [PS2] to the compounding flywheel in [PS1]:

  1. Acting as the Transfer Gate for [PS2]: In [PS2], teams explore new ideas to answer a simple question: can we make this work? But before an idea asks for real company capital, it must cross the testing gate. Hopkins' testing engine answers the financial question: does this pay for itself? Plant managers do not have to take green claims on faith; they have the operational evidence needed to justify the investment.

  2. Supplying Verified Baselines for [PS3]: In [PS3], we set up OKRs to drive quarterly execution. An effective Key Result requires an accurate, validated baseline. Controlled test campaigns generate the exact baseline data needed to set ambitious, realistic quarterly performance targets.

  3. Protecting the Compounding Loop for [PS1]: By enforcing small-scale testing and control groups, you prevent capital from flowing into low-yield pet projects. Only initiatives with proven unit economics clear the gate into the Better decisions → Savings & innovation stage. This discipline ensures that projected savings turn into real cash that flows directly into the reinvestment cycle.


Case study – Opower: randomized trials before national scale

A modern demonstration of disciplined testing in sustainability comes from energy software company Opower.

User's avatar

Continue reading this post for free, courtesy of Lars Wullink.

Or purchase a paid subscription.
© 2026 Sustainability Simplified · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture