The decision, the way the market tells it
In 2009 the Empire State Building announced a $106 million energy retrofit. The headline traveled the world: 38% less energy, $4.4 million saved every year, a three-year payback. Fifteen years later it is still the case study wheeled into boardrooms to justify a deep-retrofit budget.
If you own a large building, you have almost certainly seen a version of this number in a vendor's deck: “look what the Empire State Building did, this pays for itself.” We didn't set out to catch anyone. We ran the decision the way the market actually uses it. Can this retrofit be relied on as proof that deep-retrofit capital delivers? We put that question through a framework that does one thing: it refuses to let a decision advance faster than its evidence allows, and it names precisely what evidence would let it. Here is what came back, using nothing but public records.
Who moves the number matters as much as the number.
The one question that governs the decision
Every number above answers the same implicit question: is this building inefficient, and will these eight measures fix it? The question that actually governs a building like this is the one the vendor's spreadsheet never asks:
Does the owner actually control the load being measured, and is the change we're seeing caused by the retrofit, or by something else entirely?
In plain terms: in a 2.7-million-square-foot tower full of tenants who control their own thermostats and pay their own meters, who moves the number matters as much as the number. And a savings claim is only worth what its baseline is worth, the “before” you measure against. Hold that question. Everything below is what the public record can, and cannot, say about it.
What the public record actually shows
New York City has required this building to report its measured energy use every year since 2012, under Local Law 84. That is not a model. It is the meter.
Site EUI (kBtu/ft²/yr) is the building's “miles per gallon.” Lower is better. Source: NYC Local Law 84 disclosure, BIN 1015862. The dashed line is the retrofit's projected target of 60; the shaded band marks the pandemic years, when the building was far from full.
| Year | Site EUI | Weather-normalized¹ | ENERGY STAR² |
|---|---|---|---|
| 2012 | 74.6 | 81.4 | 84 |
| 2013 · retrofit done | 81.3 | 80.1 | 82 |
| 2014 | 82.2 | 79.8 | 83 |
| 2015 | 83.1 | 82.5 | 84 |
| 2016 | 72.6 | — | 87 |
| 2017 | 75.0 | 78.3 | 86 |
| 2018 | 80.7 | 80.2 | 77 |
| 2019 | 80.6 | — | 80 |
| 2020 · pandemic | 58.9 | — | 81 |
| 2021 · pandemic | 56.0 | — | 79 |
| 2022 | 64.4 | 65.1 | 80 |
| 2023 | 59.4 | 63.5 | 79 |
¹ Weather-normalized: the city's own correction for hot and cold years, so you compare the building to itself, not to the weather. ² ENERGY STAR (1–100): how this building ranks against similar buildings nationwide, a percentile, not a measurement of savings.
Read carefully, and stated no more strongly than the record allows: on the city's own weather-normalized measure the building sat flat between 78 and 83 from 2012 through 2018, six years with no downward trend, while the goal was 60. It does operate lower now: 65.1 in 2022 and 63.5 in 2023. That improvement is real, weather-corrected, and sustained. But its timing does not match the retrofit. The work finished in 2013; the step down appears around 2020, seven years later, and the public record does not identify what caused it.
There is a real, measured improvement. There is a famous project it is attributed to. And the dates don't line up. Public data can establish the gap. It cannot close it, because closing it needs evidence the public record doesn't hold.
A real improvement, arriving seven years after the work it's credited to.
The famous payback is a choice of denominator
The “three-year payback” isn't wrong. It is a choice of what you put on the bottom of the fraction.
| What you count as “capital” | ÷ $4.4M/yr |
|---|---|
| $13.2M, the incremental cost | ≈ 3.0 years |
| ~$92M, the eight measures' own stated costs | ≈ 20.9 years |
| $106M, the energy-project budget | ≈ 24.1 years |
The three-year number treats about $79M of the work (new air handlers, tenant lighting) as “we were going to spend it anyway.” Maybe true, but that is an assumption about a counterfactual nobody measured: the business-as-usual budget that would have existed without the energy goal. Change that one assumption and a 3-year payback becomes a 24-year payback.
One more tell. The 2010 white paper's section titled “Energy Use Operating Data” contains, in its own words, modeled data, not operating data. An independent review by Lucas Davis (UC Berkeley, Haas School of Business) made the same point: most of the public reporting describes simulated consumption, and real measured performance was what was needed to validate it. For context, the national median site EUI for a building of this type is about 96.5, so the tower is genuinely efficient against its peers. That is a different claim from “the eight measures delivered 38%.”
Why the honest answer is “not yet”
Most analysis stops here with a shrug, “we can't really know without more data.” A governed read does the opposite. It states exactly how much it knows, on a fixed ladder, and exactly what would move it up.
| Level | What it means |
|---|---|
| L0 | we haven't observed it |
| L1 | someone told us / a benchmark |
| L2 | typical for this kind of building |
| L3 | documented for this building (measured EUI, emissions) |
| L4 | independently verified / metered |
The building's energy number reaches L3, it's the city's mandated measurement. But the decision doesn't run on the number alone. It runs on a cluster of variables, and here the read applies a rule any engineer will recognize:
A chain is only as strong as its weakest link. A cluster can't be graded higher than its weakest member.
The “fuel & energy” cluster needs four things: the measured EUI ✓ (L3), the emissions ✓ (L3), the utility bills ✗ (L0), and the primary-fuel / control picture ✗ (L0). Two of four are documented; two are not seen, so the cluster, and the decision, are held at the floor. The verdict isn't “no.” It is “not on this evidence, and here is the evidence that changes it.”
The evidence that unblocks the case is already in a compliance folder.
New York already made you produce the evidence that lifts this
Here is the quietly remarkable part. The measurement that blocksthe case and the evidence that unblocks it were both produced by the same owner, by law. The city's energy laws aren't a list, they're a ladder of evidence, and the rungs you're missing are sitting in a compliance folder.
| The law | What it made you produce | What it lifts |
|---|---|---|
| LL84 · Benchmarking | your annual energy number | already used → L3 |
| LL87 · Energy audit + retro-commissioning | a systems-level audit of where the energy goes | the systems picture |
| LL88 · Sub-metering + tenant statements | who controls and pays for each load | the control boundary, the governing variable |
| LL97 · Emissions cap + penalty | the stakes, and your filed basis | the regulatory picture |
LL87: buildings over 50,000 ft² must complete an energy audit and retro-commissioning every 10 years, and that audit is the “why” behind your EUI. LL88: sub-meters in tenant spaces over 5,000 ft², with monthly statements, almost exactly the control-boundary evidence the read is asking for. LL97: penalty = (reported emissions − your cap) × $268 per ton over the limit, the number everyone is deciding against. (Related rungs: LL85 folds the energy code into renovations; LL95 turns your LL84 score into the letter grade on the front door; LL92/94 govern new roofs.)
What your own evidence does
| You add | The verdict does this |
|---|---|
| your LL88 sub-metering | the governing question gets an answer; deferred lanes open |
| your LL87 audit | the “where the energy goes” gap closes |
| 12 months of utility bills | the fuel-energy cluster clears its weakest link |
| a bill-backed baseline | the withheld savings / payback figures become computable |
Same building, same framework. The difference isn't a better model, it's your own paperwork. The read gets more precise as you feed it, not more vague. And it never fills a blank for you: if you don't have a document, it stays a blank, named, on the list.
The honest verdict
On public evidence alone, three things hold, and no more. The improvement is real, but its cause is unproven (it arrives ~7 years after the retrofit it's credited to). The famous payback is a choice of denominator, 3 to 24 years on the same invoices. And the public record cannot confirm the savings claim, but it names precisely which of your documents would. None of this says the retrofit was bad, or that anyone acted in bad faith. It says something more useful: a decision this size deserves to know which of its numbers are load-bearing and which are assumptions, before the committee votes, not after.
Method: this is a governed read produced by the źlab framework on public sources only. It states what the evidence permits and prohibits asserting; it does not recommend actions. Where the record is silent, it says so. The framework's report ceiling for this case is a Decision-Blocked Asset Brief, held for site evidence, by design.
Sources: NYC Open Data · Energy & Water Data Disclosure (Local Law 84), BIN 1015862 · Empire State Realty Trust SEC filings · Jones Lang LaSalle, “One building can change the world” (2010) · Rocky Mountain Institute case study (2009) · L. Davis, “Ex Post Evaluation of the Empire State Building Retrofit,” Energy Institute at Haas.
Your asset could be the next read.
A governed read of your own decision, built with the same discipline as this one, and it gets sharper the moment you add the evidence you already hold.