The Tomography of a City: Reading Airlines, Rental Cars, Grocers, and Gyms to Find Growth Before It Prints
A friend of mine runs an industrial fund. Six or seven years ago he told me, in the tone of someone confessing to a mild vice, that he had started reading airline schedules.
Not flying. Reading. He had a spreadsheet of Cirium schedule data and he checked it the way other people check the ten-year. His argument was that when an airline commits an aircraft to a new nonstop, it has made a specific, expensive, forward-dated judgment about the volume of people who need to be in two places, and that judgment is published eleven months before the first flight, for anyone who cares to look. Few people in real estate looks.
He was buying shallow-bay industrial in a market whose only apparent virtue at the time was that it was cheap. Then a carrier put a daily nonstop on it to a coastal tech hub, on a mainline narrowbody rather than a regional jet, and kept it through the winter. He bought two more buildings. That trade worked out extremely well and he still talks about the schedule filing, not the buildings.
I have thought about that conversation for years, and this paper is what came out of it.
The premise is simple to state and takes some work to use. Dozens of very sophisticated companies are running expensive location models on your submarket right now, and they publish the results. Not the models. The conclusions, in the form of land purchases, permit applications, lease signings, schedule filings, and franchise disclosures. Each one is a dated, dollar-denominated, capital-backed forecast made by people whose careers depend on being right. You can read all of them for free.
The catch, and it is a serious one, is that they are not answering your question, they are not independent of each other, and the naive way of combining them will make you dramatically more confident than the evidence supports. Most of this paper is about that catch.
On sourcing. Site selection criteria for private companies are not published. The thresholds cited here come from executive interviews, investor presentations, trade press, permit records, and reverse-engineering from observed store locations. They are directionally reliable and should not be treated as official. Where a number is well documented I say so. Where it is folklore that happens to be useful, I say that too.
Part I: Somebody Already Did Your Market Study
Consider what happens before a Costco opens.
A team spends eighteen to thirty-six months on the site. They model the trade area, typically something like 150,000 to 250,000 people within a ten to fifteen mile drive, adjusted for competition, income distribution, and drive-time geometry rather than radius. They project membership penetration and renewal, which is the part that matters most, because the membership model means Costco is not betting on people shopping there, it is betting on people staying there and paying again next year. Then the company buys the land. Fifteen to twenty acres, usually in fee, frequently at the top of the local market, and builds a $40 million to $60 million box on it.
That is a thirty-year forecast with the balance sheet behind it. Nobody does that on instinct.
And when the site plan hits the county planning department, that entire forecast becomes a public document. Not the analysis, but the conclusion, which is the part with the money on it. You are being handed the output of a research budget you could not afford, by an organization with no interest in whether you find it useful, in a filing that costs you nothing to read.
Now run the same logic across the whole set. Delta commits an airframe. Enterprise signs a lease on a neighborhood branch. Trader Joe's takes 13,000 square feet on a fifteen-year term. Life Time files for a $30 million athletic club on eight acres. Every one of these is a forecast with a date, a dollar amount, and a name attached.
The industry treats these as amenities. "The site benefits from proximity to a new Costco." That sentence appears in a thousand offering memoranda a year and it is doing almost no work. It is being used as a comfort signal, roughly equivalent to mentioning that the property has good schools. What it should be used as is evidence about the future, weighted by how much the person making the claim had to risk in order to make it.
So the first reframe is this: stop reading anchors as amenities and start reading them as forecasts. An anchor is not a feature of your site. It is somebody else's opinion about your site, expressed in the only language that cannot be faked, which is capital they cannot easily get back.
Part II: Everyone Is Looking at the Same City With a Different Instrument
Here is the part that took me longest to understand, and it is the organizing idea of the whole method.
Every one of these companies needs to know the same underlying thing. Call it the latent variable: how many of what kind of households will be here, with how much money, behaving in what way, for how long. That is the quantity that determines whether a warehouse club works, whether a nonstop fills, whether a gym hits membership targets, and whether your apartment building leases at pro forma.
Nobody can measure it directly. It does not exist as a number anywhere. So each industry measures a projection of it onto whatever axis its own economics happen to care about.
| Player | Thinks it is measuring | Is actually measuring | Projection of |
|---|---|---|---|
| Airline network planner | Route contribution margin | Corporate density, travel budgets, second-home wealth | Economic base composition |
| Rental car fleet manager | Fleet utilization and turn | Ground-level business activity, car dependency, accident volume | Settlement pattern and employment intensity |
| Grocer | Trade area sales capture | Household count times income times education, net of competition | Household formation |
| Warehouse club | Membership renewal rate | Household stability, family stage, car ownership | Residential permanence |
| Gym operator | Membership capture and churn | Discretionary income, life stage, commute geometry | Household affluence and stability |
| Apartment developer | Rent and absorption | All of the above | All of the above |
Read the last row against the others. The real estate developer is the only participant who needs the entire latent variable, and the only one collecting none of the projections.
The right mental model for what to do about that is a CT scanner.
A CT scanner cannot see inside you. What it does is take a large number of one-dimensional X-ray projections from many different angles and reconstruct the three-dimensional object mathematically. Any single projection is nearly useless; a chest X-ray from one angle collapses your entire torso onto a plane and throws away depth. But projections from enough distinct angles contain, between them, everything.
That is exactly what is available here. The airline's view of a metro is one projection. The grocer's is another, taken from a completely different angle, because the grocer cares about household count and the airline does not care about households at all. The gym's is a third. None of them is the market. Together, if you have enough of them and they come from sufficiently different angles, they reconstruct it.
Hold onto the qualifier, because Part V is about what happens when you violate it. In tomography, if all your projections are taken from nearly the same angle, you cannot reconstruct anything. You get a characteristic smear. Radiologists call it limited-angle artifact, and the dangerous property is that the reconstruction still looks like an image. It just is not one.
Before that, the instruments.
Part III: The Instruments, One at a Time
IIIa. Airlines: the only genuinely forward-published signal in the set
Airlines are the most underused data source in commercial real estate and I do not think it is close.
What the decision actually optimizes. An airline network planner is solving a fleet assignment problem. They have a finite number of airframes, each of which must fly a connected sequence of legs that returns it to a maintenance base, staffed by crews with legal duty limits, using gates they have leases on. Route profitability matters, but it enters as a constraint on a much larger optimization about aircraft utilization. This is the crucial fact for our purposes: the airline is not thinking about your submarket at all. Any demographic information in its decision is a byproduct of solving an unrelated problem, which, as we will see in Part VI, is precisely what makes the signal valuable.
What is public. More than most people realize.
T-100 (Bureau of Transportation Statistics) gives monthly segment and market data by carrier: passengers, seats, load factor, freight. Free, roughly a three-month lag.
DB1B was the ten percent quarterly sample of tickets that included fare and true origin and destination, and it is the source most analysts still cite. Note that it was discontinued as of July 2025 and replaced by DB1C, a forty percent sample collected monthly. If you are building anything on this, build it on DB1C. Four times the sample at three times the frequency is a materially better instrument, and almost nobody has updated their pipelines.
Schedules (OAG, Cirium) are the forward-looking piece and the reason to care. Carriers file schedules six to eleven months ahead. This is the only publicly available dataset in this entire paper where somebody tells you what they are going to do before they do it.
The tells that are real.
Gauge, not frequency. Adding a second daily flight can be an artifact of where an airframe needed to be overnight. Swapping a 76-seat regional jet for a 190-seat narrowbody on the same route is a 150 percent capacity increase that required someone to argue for it in a meeting. Upgauging is a much stronger statement than frequency.
Point-to-point beats hub-feed. This is the distinction that separates people who use this data from people who quote it. A new nonstop into a carrier's own hub tells you almost nothing, because the carrier needs to feed the hub and the route economics are network economics. A nonstop between two cities where neither is that carrier's hub is a pure judgment about origin and destination demand. Somebody looked at the number of people who need to get from A to B and concluded there were enough of them to fill an airplane every day without any connecting traffic. Non-hub point-to-point route additions are the single highest-quality metro-level growth signal available for free.
Survive the second winter. Airlines launch a great many routes and quietly kill a great many. A route that makes it through two winters has cleared the seasonality test and is probably real.
Yield rising with capacity rising. This is the best one and it requires DB1C. In a normal market, adding seats pushes fares down. When capacity is rising and average fare is also rising, demand is outrunning supply, which in a domestic market almost always means business travel, which means corporate presence, which means office absorption twelve to eighteen months out. There is no demographic dataset in existence that contains this information.
Booking curve and day-of-week shape. A market with sharp Monday morning and Thursday evening peaks and short advance purchase windows is a business market. A market with Friday through Sunday peaks and ninety-day booking curves is a leisure market. These predict entirely different real estate. One says office and extended-stay. The other says hospitality, short-term rental, and second homes. You can distinguish them from a schedule and a fare sample, and I have never once seen this done in a market study.
The mistake people make. Treating passenger counts as the signal. Enplanements are a lagging measure of what already happened and are heavily contaminated by connecting traffic at hub airports. The signal is in the forward schedule and in the fare, not in the headcount.
IIIb. Rental cars: the migration tape nobody reads
What the decision optimizes. Two almost unrelated businesses wear the same brand. The airport business is a concession operation: bid for space, pay a minimum annual guarantee, optimize fleet turn against arriving passengers. The neighborhood business is mostly insurance replacement, which is to say it exists because cars get damaged and insurers need to put people in loaners. Enterprise is the interesting operator here, with roughly 4,100 to 4,700 US locations and a stated position within fifteen miles of ninety percent of the American population.
The tells.
Neighborhood branch openings are a proxy for car-dependent household density. Insurance replacement volume is a function of vehicles on the road, accident rates, and body shop capacity. A new neighborhood branch in a growing suburb is a small, cheap, quiet vote that there are now enough households with enough cars. Low conviction individually, useful in aggregate, and roughly independent of everything else in this paper.
One-way pricing asymmetry. This is my favorite signal in the whole set because it is a live, market-priced migration indicator sitting on a public website.
Rental fleets have to be rebalanced. If more people rent one-way from Chicago to Austin than the reverse, the fleet accumulates in Austin and drains from Chicago, and the company must either pay to truck cars back or price the imbalance away. So they price it away. Check the same one-way rental, same dates, in both directions. If Chicago to Austin is $600 and Austin to Chicago is $89, that spread is a market-clearing price on net directional flow.
U-Haul has built a whole growth index on exactly this logic and gets a news cycle out of it every January. The rental car version is more granular, updates daily, covers business relocations as well as household moves, and requires nothing but a browser and some patience. Sample it weekly across thirty city pairs for a year and you have a proprietary migration dataset built from public prices.
Airport concession bids. When a rental company bids for airport counter and garage space, it commits to a minimum annual guarantee over a multi-year term. Those bids and MAGs sit in airport authority procurement records, which are public. A rising MAG is a contractual forecast of passenger volume, underwritten by somebody with money at stake. This is obscure enough that I have never heard it discussed, and it is sitting in board packets at every mid-sized airport authority in the country.
The mistake. Reading the airport business and the neighborhood business as the same signal. They measure different things and can move in opposite directions in the same metro, which is itself informative: strong airport, weak neighborhood means a market attracting visitors and not residents.
IIIc. Grocery and big box: the highest conviction, the latest arrival
Trader Joe's. The founder, Joe Coulombe, said he would not consider a trade area with fewer than about 40,000 households likely to contain his core customer. The company targets college-educated households, commonly cited at median income above $100,000, in compact areas where people live and shop within a tight radius. Store footprint runs roughly 12,000 to 15,000 square feet, no pharmacy, no service deli, and famously indifferent parking.
That small footprint is why Trader Joe's is the best mid-stage urban signal available. It can physically fit into places a full-line grocer cannot, so it arrives earlier. And its screen weights education and density more heavily than raw income, which happens to be an excellent description of a neighborhood two to four years before it reprices. If I could have exactly one retail signal for an infill submarket, this would be it.
Whole Foods, and the thing everyone gets wrong. Before 2017, a Whole Foods was a straightforward affluence signal. After Amazon acquired it, store siting became partly a logistics decision, because the stores function as nodes in a delivery and pickup network. A Whole Foods now tells you something about Amazon's density and last-mile model as well as about local incomes. The signal did not stop working. It changed meaning, and most people are still reading the pre-2017 version.
Costco. The heaviest commitment in retail and therefore the highest conviction and the latest. Fifteen to twenty acres, fee ownership, a trade area typically cited at 150,000 to 250,000 people within a ten to fifteen mile drive, skewed toward middle and upper-middle incomes. The membership model means Costco is underwriting renewal, which makes it a bet on household stability rather than household count.
Here is the part that matters and gets missed: for a land buyer, Costco is a confirmation, not a prediction. By the time the site plan is filed, the land trade is over. But for a merchant builder of apartments, the same signal is still early, because Costco's opening precedes by three to five years the household growth it forecast. One signal, two completely different lead times, depending on which asset class you are underwriting. Any framework that assigns a single "lead time" to a signal is broken.
Target, two companies in one. A 130,000 square foot full-format store is a suburban household formation signal with a five to ten mile trade area. A 20,000 square foot small format store is an urban density signal with a walkshed. Same logo, opposite meanings. Coding them the same way in a dataset is a common and expensive error.
Three tells that beat all of the above.
The second store. Everybody tracks "Costco is coming to X." Almost nobody tracks "Costco is opening a second warehouse seven miles from the first." That is a far stronger signal, because the company has concluded the trade area can absorb deliberate cannibalization, which means the first store is capacity-constrained, which means the trade area outperformed the original model. Infill beats entry. Entry is a forecast. Infill is a forecast that already came true and is being extended.
The renewal. New stores are subsidized. They come with tax increment financing, abatements, infrastructure contributions, and free land. A renewal comes with none of that. When a tenant re-ups a fifteen-year lease at market rent with no incentive package, that is the cleanest, least contaminated signal in the entire dataset, and it is discoverable in REIT supplemental disclosures. Nobody tracks renewals. Everybody tracks openings. The renewals are better.
The parking ratio in the site plan. If a chain whose prototype calls for five spaces per thousand square feet files a plan at 3.5 per thousand, it has made an explicit forecast that walk-up, transit, and delivery will replace the difference. That is a density prediction, stated in a public document, in a number nobody reads. Read the site plan, not the press release.
IIId. Fitness: the fastest signal and the slowest, in the same category
Fitness is useful precisely because the category spans an enormous range of capital intensity, which gives you a natural fast indicator and slow indicator pair.
Life Time, at the slow end. Prototypes run roughly 86,000 square feet in warm climates and around 102,000 in cold, on sites from about five to fifteen acres, at build costs recently permitted around $30 million and up. Member median household income is disclosed in investor materials at roughly $159,000. A Life Time is a twenty-year bet that a specific ten to fifteen minute drive-time ring will contain and retain affluent families. In terms of income selectivity per square foot of commitment, it may be the most demanding site decision in national retail.
Equinox and urban luxury. One to two mile walkshed, young affluent, correlates tightly with Class A multifamily. Smaller commitment than Life Time, faster, more reversible, and a good confirmer for a luxury rental thesis.
Boutique franchise, at the fast end. Club Pilates, Orangetheory, F45, Barry's. Three to five thousand square feet, five to seven year leases, a few hundred thousand dollars of buildout, usually a franchisee. These arrive early and they are noisy, because franchisees are far less analytically disciplined than corporate site selection teams. A meaningful share of franchise locations exist because the franchisee lives twelve minutes away.
The way to use them is as a cluster signal with a threshold, not individually. One Club Pilates is noise. Four independent boutique studios opening within eighteen months inside a one mile radius is a real statement about the density of a specific demographic, because four separate operators each independently underwrote the same walkshed.
Planet Fitness measures something else entirely. Its model wants population density and is close to indifferent to income, sometimes inversely related to it. This gives you a diagnostic: a submarket receiving Planet Fitness and Club Pilates simultaneously is bifurcating. That is genuinely useful, because it means the product strategy should be barbell. Build for one end or the other. The middle of that market is about to be the worst place to be.
The obscure source that makes this category work. Franchisors must file a Franchise Disclosure Document, and Item 20 requires state-by-state disclosure of outlets at the start and end of each of the last three fiscal years, openings, closures, terminations, transfers, and projected openings for the coming year.
Read that list again. Closures. Transfers. Projected openings by state. Legally required, standardized, free, and forward-looking.
Transfers are the sleeper. A franchisee selling their outlet to another franchisee is not a closure, so it does not show up in any count of store openings and closings, but a rising transfer rate in a state is franchisees getting out, and it reliably leads closures. It is a distress indicator in a dataset nobody in real estate opens.
Part IV: Signals Arrive in an Order
Signals do not fire simultaneously. They come in a rough sequence, and the sequence is more informative than any individual signal, because deviation from the expected order is where the real information lives.
For an urbanizing infill submarket, the canonical order runs something like:
- Independent coffee and a chef-driven restaurant. Very low capital, very fast, very noisy.
- Liquor license applications rising. Public, precedes openings by six to twelve months, and a good early-warning tape.
- Boutique fitness cluster crossing three or four operators.
- Trader Joe's.
- Small-format Target or a full-line urban grocer.
- Class A multifamily starts from a national merchant builder.
- A national office tenant signing a full floor.
For a suburban growth corridor:
- Residential building permits inflecting.
- Dollar General or a value grocer. These arrive very early and are often dismissed, wrongly.
- Enterprise neighborhood branch.
- Full-format Target.
- Costco.
- Life Time.
- A hospital system outpatient campus, which is the last and most conservative actor in the set.
Now the useful part. Out-of-order signals are the trade.
A submarket that gets a Trader Joe's without the boutique fitness cluster that normally precedes it is either an early-stage market where Trader Joe's is being unusually aggressive, or a market with a demographic anomaly, most often a university or a hospital campus producing educated households without the discretionary spending pattern. Worth a phone call either way.
A submarket that gets a Costco without the residential permit inflection that normally precedes it is usually a geographic decision rather than a demographic one. Costco is filling a hole in its drive-time coverage map, capturing existing households from an adjacent overloaded store. That is a real business decision and it is not a growth forecast, and if you underwrite it as one you are going to be disappointed.
And a submarket where the fast signals fire but the slow ones never follow, where the boutique studios and the restaurants show up but no anchor ever commits, is telling you something specific and bad: the demographic is there and the depth is not. This is the classic profile of a submarket that supports one good block and never becomes a district.
Part V: The Circle
Everything to this point is the optimistic version. Now the problem, which I think is the most important idea in this paper and which invalidates most of how alternative data is currently used in site selection.
The loop
Here is what actually happens, step by step.
A handful of location intelligence vendors aggregate mobile device panels, credit card panels, and census-derived demographic estimates. Call them what they are: Placer, Esri, Claritas, Buxton, and a few others. They sell substantially the same underlying panels to Costco, to Target, to Life Time, to Whole Foods, and to the brokerage advising the apartment developer down the street.
Each of those teams runs a structurally similar gravity model on substantially the same data.
They reach similar conclusions. They site near each other. Part of that is genuine co-tenancy economics, and part of it, the part nobody says out loud, is that they are running correlated models on correlated inputs.
The resulting cluster of anchors causes residential developers to underwrite the submarket, because "anchored by Costco and Target" is the second bullet on every offering memorandum ever written.
The apartments get built. The households arrive.
The mobile panel registers the households. The vendor data now shows the trade area growing exactly as forecast.
Every model in the system gets more confident.
There is no step in that loop where an independent observer checks whether the original forecast was correct. The forecast was made true by the act of making it. Capital followed the signal, and the capital produced the outcome the signal predicted.
And the loop runs in reverse with equal force. A submarket that narrowly misses the anchor threshold never receives the residential capital that would have carried it across the threshold. It then fails to grow, which confirms the original rejection. The threshold is not a property of the submarket. It is partly a property of the decision.
Which leads somewhere slightly uncomfortable. A meaningful share of American urban growth is being allocated by a few hundred site selection professionals at perhaps fifty companies, working from four data vendors, none of whom think of themselves as urban planners and all of whom would reject the description.
The arithmetic of false triangulation
Here is what the loop does to your confidence, and this is the number I would put on the first slide.
Suppose you are evaluating a submarket. Your prior that it materially outperforms the metro is 30 percent. You then collect five signals: a Costco land purchase, a Trader Joe's opening, a Life Time site plan, a boutique fitness cluster, and a Target announcement. Assume each signal is 70 percent accurate, both sensitivity and specificity, which is generous.
Each signal carries a likelihood ratio of 0.7 / 0.3 = 2.33. Prior odds are 0.3 / 0.7 = 0.4286.
If the five signals are independent:
Posterior odds = 0.4286 × 2.33⁵ = 0.4286 × 69.16 = 29.64
Posterior probability = 96.7 percent.
That is the number the naive triangulation produces, and it feels right. Five separate sophisticated companies all committed capital to this submarket. How could they all be wrong.
Now assume the signals are correlated at ρ = 0.6, which is conservative given that all five teams bought overlapping panels from overlapping vendors. Effective independent sample size is n / (1 + (n−1)ρ) = 5 / 3.4 = 1.47 signals.
Posterior odds = 0.4286 × 2.33^1.47 = 0.4286 × 3.48 = 1.49
Posterior probability = 59.8 percent.
And if the five signals are actually one signal echoed five times, which is the limiting case where every team ran the same model on the same panel:
Posterior probability = 50.0 percent.
A coin flip.
Same five confirmations. Same underlying evidence. Ninety-seven percent, sixty percent, or fifty percent, and the entire difference is a correlation coefficient that nobody in this industry estimates, reports, or thinks about.
Push it one step further and ask what independence you would actually need to justify a 90 percent conviction. Solving backward, you need about 3.6 effectively independent signals out of your five, which requires correlation below roughly 0.10.
Five retail site models built on shared vendor panels are not correlated at 0.10. They are not close. Which means that in the ordinary case, the honest posterior after five glowing retail confirmations is somewhere in the fifties or sixties, and the ninety-seven is an artifact of an assumption nobody made explicitly.
This has already happened once
Between roughly 2003 and 2007, essentially every retail site model in America pointed at the exurban power center. Household formation was moving outward, the demographic projections were sound, and the models agreed.
They built. And the building itself validated the models, because each new center appeared in the others' competitive data as evidence of a thickening trade area. Between 2008 and 2012 the country worked through several hundred functionally dead centers.
The models were not wrong about the demographics. That is the part worth sitting with. Household formation really was moving outward, and it resumed doing so afterward. The models were wrong about their own independence. Each retailer believed it was confirming an external fact about the world. What it was actually confirming was the other retailers.
That is not a demographic failure. It is an epistemics failure, and the machinery that produced it is more tightly coupled now than it was then, because in 2005 the models at least used different data.
Part VI: Breaking the Circle
If correlation is the disease, independence is the cure, and there is a clean principle for finding it.
Prefer signals generated as a byproduct of an optimization that has nothing to do with real estate.
An airline network planner is solving fleet assignment. A rental car fleet manager is solving vehicle repositioning. A state alcoholic beverage board is processing license applications. None of them is forecasting your submarket, none of them buys the location intelligence panel, and none of them is inside the loop. Whatever demographic information falls out of their decisions is uncontaminated by the reflexive machinery described above.
That single criterion reorders the entire signal set.
| Signal | Independence | Why |
|---|---|---|
| Airline yield and capacity from DB1C | 0.95 | Fleet economics, entirely outside the loop |
| One-way rental car price asymmetry | 0.95 | Fleet rebalancing, market-priced, outside the loop |
| New non-hub point-to-point route | 0.90 | Network planning, outside the loop |
| Liquor license applications | 0.85 | Regulatory queue, independent process |
| IRS county-to-county migration with AGI | 0.85 | Tax filings, no commercial panel involved |
| Enterprise neighborhood branch | 0.80 | Insurance replacement economics |
| School enrollment by grade cohort | 0.80 | Administrative data, and grade cohorts reveal family stage |
| Boutique fitness cluster | 0.60 | Franchisees, semi-independent judgment, noisy |
| Second-store infill by an anchor | 0.60 | Same models, but validated by actual observed sales |
| Trader Joe's, Costco, Life Time | 0.50 | Sophisticated, and squarely inside the vendor loop |
| Target, Whole Foods | 0.40 | Deeply inside the loop, extensive shared vendor use |
| Placer-derived foot traffic rankings | 0.15 | This is the loop |
Look at what happens to the ranking. The signals with the most capital behind them are the ones you should trust least per unit of confirmation, because they are the most correlated with each other. The signals with the least capital behind them, if they come from outside the system, carry more independent information.
That is deeply counterintuitive and it is the central practical finding here. Conviction and independence trade off, and the industry optimizes entirely for conviction.
A useful heuristic falls out of it: one airline yield signal is worth more than three retail confirmations, because the three retail confirmations are mostly the same confirmation.
Part VII: Scoring a Submarket
Frameworks that do not produce a number are opinions. Here is the number.
Weight each signal by conviction times independence. Conviction is a zero to ten score reflecting capital at risk and irreversibility. Independence is the multiplier from the table above.
| Signal | Conviction | Independence | Weight |
|---|---|---|---|
| Yield rising with capacity rising | 7 | 0.95 | 6.65 |
| Second-store infill by an anchor | 9 | 0.60 | 5.40 |
| New non-hub point-to-point route | 6 | 0.90 | 5.40 |
| Costco land acquisition | 10 | 0.50 | 5.00 |
| Sustained one-way rental asymmetry | 5 | 0.95 | 4.75 |
| Route upgauge | 5 | 0.90 | 4.50 |
| Life Time or equivalent big-box fitness | 9 | 0.50 | 4.50 |
| Trader Joe's | 7 | 0.50 | 3.50 |
| Target full format | 8 | 0.40 | 3.20 |
| Enterprise neighborhood branch | 4 | 0.80 | 3.20 |
| Boutique fitness cluster, three or more | 5 | 0.60 | 3.00 |
| Whole Foods | 7 | 0.40 | 2.80 |
| Target small format | 6 | 0.40 | 2.40 |
Note the top of that table. The highest-weighted signal is an airline fare statistic, and the second is a retail event almost nobody tracks. Costco, the most capital-intensive decision in the set, ranks fourth.
Metro-level signals get a 0.5 spatial multiplier when applied to a specific submarket, because a nonstop is a fact about a region and you are underwriting eight blocks.
Two submarkets
Submarket A, "the Corridor." Exurban, twenty-two miles out, greenfield, farmland converting to subdivisions.
| Signal present | Weight |
|---|---|
| Costco land acquisition, closed 14 months ago | 5.00 |
| Life Time site plan filed | 4.50 |
| Target full format announced | 3.20 |
| Enterprise neighborhood branch, 8 months old | 3.20 |
| Metro: new non-hub route plus two upgauges (9.90 × 0.5) | 4.95 |
| Raw score | 20.85 |
Submarket B, "the Node." Infill, three miles from the CBD, warehouse district converting.
| Signal present | Weight |
|---|---|
| Trader Joe's, opened 20 months ago | 3.50 |
| Boutique fitness cluster, four studios in 16 months | 3.00 |
| Target small format | 2.40 |
| Whole Foods, 12 years old, decayed to 0.3 of value | 0.84 |
| Metro: same as A | 4.95 |
| Raw score | 14.69 |
Submarket A wins by 42 percent. And that comparison is completely invalid.
The correction that makes it work
Submarket B did not fail to attract a Costco. Submarket B is physically incapable of attracting a Costco, because Costco needs fifteen to twenty acres and there is no such parcel within three miles of a downtown. The same is true of Life Time.
The absence of those signals in B is information about parcel geometry, not about demand. Scoring B against a ruler that includes them is like grading a swimmer on their vertical jump.
So normalize. Score each submarket as a percentage of the signals it is structurally capable of emitting.
Attainable for A, a greenfield corridor with unlimited large parcels: Costco 5.00 + second-store 5.40 + Life Time 4.50 + Trader Joe's 3.50 + Target full 3.20 + Enterprise 3.20 + Whole Foods 2.80 + metro 4.95 = 32.55
Attainable for B, an infill node with no large parcels: second-store 5.40 + metro 4.95 + Trader Joe's 3.50 + Enterprise 3.20 + boutique cluster 3.00 + Whole Foods 2.80 + Target small 2.40 = 25.25
| Raw | Attainable | Normalized | |
|---|---|---|---|
| Submarket A | 20.85 | 32.55 | 64% |
| Submarket B | 14.69 | 25.25 | 58% |
A 42 percent raw advantage collapses to a 10 percent normalized advantage. These are two comparably strong submarkets, and the raw score was measuring parcel size.
That normalization step is, in my experience, the difference between this method working and this method systematically routing capital to the suburbs for reasons that have nothing to do with growth.
And the score is not the answer anyway
Both submarkets score in the high fifties to mid sixties. Neither number tells you what to build. The composition does.
A's signals are all family-formation and car-dependency signals: warehouse club, athletic country club, full-format discount, replacement rentals. That is a for-sale single family and garden apartment thesis with a large-unit mix and generous parking.
B's signals are all density-and-education signals from operators with small footprints and no parking requirements. That is a Class A infill multifamily thesis with a small-unit mix, structured parking at a ratio well below code, and ground floor retail.
Two similar scores, two opposite products. Anyone who reads the score and skips the composition has extracted maybe a fifth of the available information.
Part VIII: How This Fails
Six failure modes, in descending order of how often I see them.
Incentives contaminate everything. A Costco that received $18 million of tax increment financing, free site work, and a twenty-year abatement is not primarily a market signal. It is a subsidy signal, and the market component of the decision has been substantially bought off. Most states maintain incentive disclosure databases. Net out the incentive before you score the signal, and if the incentive was large enough, discard it. A subsidized anchor tells you about a county's economic development budget.
Availability bias in real estate, not demand. Chains go where a suitable site is available at an acceptable price, which is not the same as where demand is highest. A submarket with strong fundamentals and no fifteen-acre parcels will never emit a Costco signal. This is the same problem as the normalization above and it also runs the other way: a weak submarket with one perfect available site may capture an anchor its fundamentals do not justify.
Survivorship, and the invisible rejections. You observe the openings. You do not observe the two hundred sites the chain rejected, and the rejections contain more information than the acceptances, because they are more numerous and more discriminating. You can partially recover them: brokers know which LOIs died, planning departments hold site plans that were filed and withdrawn, and a chain that studied your submarket and passed is a data point worth more than most of the ones in your model. This is legwork rather than data science, which is exactly why it stays valuable.
Corporate distress masquerading as a market read. When a chain pauses expansion nationally because its balance sheet is strained or its cost of capital moved, that is a capital markets event, not a statement about your submarket. Check whether the pause is national before you read it locally. Several categories went quiet in 2023 through 2025 for reasons that had nothing to do with any trade area.
Franchise noise. Franchisee-selected sites carry meaningfully less analytical weight than corporate-selected sites. Item 20 helps here, because a location that gets transferred within three years of opening was probably a bad site chosen for a bad reason.
Lag asymmetry, and the derivative. Retail confirms what already happened. If you want to be early you need the rate of change, not the level. Three anchors in eighteen months is a different market from three anchors over nine years, and a static count cannot tell them apart. Every signal in your system should be timestamped, and every score should be computed as both a level and a trailing slope.
Part IX: What to Actually Build
The stack, in the order I would build it, cheapest and highest-value first.
Free and immediately useful.
- BTS T-100 for monthly segment traffic, and DB1C for fare and true origin-destination, monthly at a 40 percent sample. Update your pipeline if it still says DB1B.
- IRS Statistics of Income county-to-county migration, which uniquely includes adjusted gross income, so you learn not just how many people moved but how much money moved with them. This is the single best free migration dataset in existence and it is badly underused.
- Census Building Permits Survey, monthly, by place.
- State alcoholic beverage control license applications. A six to twelve month lead on restaurant openings.
- County planning department filings, including withdrawn ones.
- State economic development incentive disclosure databases.
- Franchise Disclosure Documents, Item 20. Openings, closures, transfers, and projected openings, state by state.
- REIT supplemental disclosures for tenant-level lease expirations and renewals.
- Airport authority board packets for concession minimum annual guarantees.
- One-way rental car quotes, sampled weekly across a fixed panel of city pairs. Thirty pairs, one script, and in twelve months you own a dataset nobody else has.
Paid, and worth it.
- OAG or Cirium forward schedules. This is the only genuinely forward-looking commercial dataset in the set and it is the one I would buy first.
- CoStar or equivalent for the real estate baseline.
Paid, and buy with your eyes open.
- Placer, Esri Business Analyst, Claritas. Useful, and this is the echo chamber. Weight anything derived from it at a fraction of what its apparent precision suggests, and never count it as an independent confirmation of a retail signal, because the retail signal was probably generated from it.
What to do first. Pick three submarkets you already know well, ideally one that worked, one that failed, and one that is live. Build the signal history for each, backdated. See whether the framework would have called them. It will be wrong somewhere, and where it is wrong is worth more than where it is right, because that is where you learn which signals your particular market actually respects. Markets differ. The weights in this paper are a starting point and should be re-estimated locally by anyone who intends to trade on them.
What This Is Really About
The instinct behind all of this is old and simple. Before there was data, good developers did this by driving around. They counted cars in parking lots on a Tuesday morning. They noticed which restaurants had a wait. They watched where the good contractors were working. They were doing exactly what this paper describes, at low resolution, using their eyes, and the best of them were extraordinarily accurate.
What has changed is not the logic. It is that the observations are now filed, timestamped, and searchable, and that the companies making them have spent enormous sums building models better than anything a developer will ever build in-house. That is a real gift and this industry mostly ignores it.
But there is a second thing that changed, and it is the reason for Part V. When every observer buys their eyes from the same four vendors, the observations stop being independent, and a method that depends on independence quietly stops working while continuing to produce numbers. The developer driving around on a Tuesday morning had terrible data and perfect independence. The modern analyst has magnificent data and almost none. Neither of those is obviously better, and the combination is much better than either.
So the discipline is this. Collect the projections, from as many genuinely different angles as you can find. Weight them by how much the observer had to risk and by how far outside your own industry's feedback loop they were standing when they made the call. Normalize for what your submarket is physically able to say. Read the composition, not just the score. And when five signals agree, before you feel good about it, ask the only question that actually determines what that agreement is worth:
Are these five observations, or is this one observation, five times?