Methodology

A short, honest account of what this model predicts, what it covers, and where its numbers come from.

What the model predicts

GridSolve models the survival of generation-interconnection queue projects — the odds that a proposed generator (solar, wind, storage, and similar) withdraws from the interconnection queue before reaching commercial operation, based on the project's attributes and queue position.

It does notmodel data center or other large-load interconnection requests. Large loads typically interconnect through a separate process at PacifiCorp and most utilities, governed by different tariffs and study procedures than the generation queue this model is trained on. GridSolve's output — deliverable capacity per substation — is a proxy for where headroom is likely to exist, not a load interconnection study.

Coverage

This dataset is national: National — 9 regions. Model validation varies by region. — about 4,687 substations across 9 ISO/RTO regions and 48 states. Withdrawal rates are computed empirically within each region, from that region's own resolved projects — they are observed history, not model extrapolation. The national, capacity-weighted withdrawal share is 61%, but regional base rates range from ~25% (ERCOT) to ~92% (NYISO).

Model reliability varies by region

The survival model's predictive quality is not uniform. It was validated separately in each region, and concordance — the probability it ranks the project that withdraws first as the higher-risk one (0.5 is a coin flip, 1.0 is perfect) — differs widely. Only some regions clear a usable bar; others were never validated at all.

RegionConcordanceTier
West0.617Marginal
Southeast0.657Usable
PJM0.597Weak
MISOUnvalidated
CAISO0.514Not predictive
ISO-NE0.651Usable
SPPUnvalidated
NYISOUnvalidated
ERCOTUnvalidated

Because the withdrawal rates are empirical, the capacity numbers stay meaningful even where the survival model is weak — the tier is a caveat about the model's forward prediction, not the observed history. Every substation's panel shows its region's tier so you can weight the score accordingly.

What predicts withdrawal

Hazard ratios above 1 raise the chance a project withdraws; below 1 lowers it. These are the features the survival model weighs most heavily.

What predicts withdrawal

hazard ratio

>1 = more likely to withdraw · caption above marks 1.0, no effect.

How the GridSolve Site Score works

Every substation with an active queue gets a single 0–100 GridSolve Site Score — our attempt to combine four separate facts into one holistic read on how suitable a site looks for a data center interconnection. It is a weighted blend of four components:

  • Scale (30%)— how much MW has queued at this substation, log-scaled against the nation's busiest sites. Heavy queue activity is evidence that real transmission-scale capacity exists here.
  • Phantom (20%) — the share of that queue expected to withdraw. Counterintuitively, high is good: a queue full of speculative projects means capacity will free up once they drop out.
  • Cost (35%) — the median historical cost to connect, log-scaled (the heaviest weight, since network upgrade cost is often the deciding factor in whether a site pencils out).
  • Viability (15%)— the inverse of this specific site's own historical withdrawal rate. High historical withdrawal is bad here: it means the site itself tends to kill projects, not just that its current queue is inflated.

Phantom versus viability is the score's subtlest distinction: a high phantom rate says this generation of the queue is speculative and will clear out; a high historical withdrawal rate says projects that try to build here tend to fail regardless. Cost is what separates a site that's cheap once the phantom projects clear (genuinely promising) from one that's just expensive everywhere (genuinely weak) — which is why cost carries the heaviest weight.

These four weights (30/20/35/15) are our editorial judgment, not something derived from the data or fit to an outcome — no outcome for "good data center site" exists in this dataset to fit against. Missing cost data is filled in with the national median and the score's confidence is reduced accordingly (shown on each substation's detail panel), and a score with few historical projects behind it is pulled back toward a neutral 50 rather than trusted at face value.

The score is also strictly relative: colors on the map and in the list are the top/middle/bottom third of the current view, recomputed live from whatever's filtered in — so filtering to one region re-bands the colors to that region. They are not fixed cutoffs and not an absolute standard.

Key limitation: this dataset has no published hosting-capacity figure, so "scale" is inferred entirely from queue activity. A substation with no current queue may be genuinely uncontested headroom, or it may simply be too small to ever serve a large load — this score cannot tell those two apart, which is also why substations with no active queue aren't scored at all rather than being guessed at.

Cost figures

Every $/kW figure on this site is a historical median network upgrade cost, drawn from prior interconnection studies at that substation. It is not a quote, not an estimate for any specific project, and can vary substantially by project size, technology, and when it enters the queue.

Queued MW

queued_mwis the MW currently sitting in the interconnection queue at a substation — it is not the utility's published hosting-capacity limit, and it is not fixed: projects enter and leave the queue continuously.

Coordinates

Verified substation coordinates are sourced from: HIFLD + Google Maps + OSM geocoding (6 teaser sites preserved) — only about 99% of substations carry them. The rest are plotted at an approximate position from their county (or, failing that, state) and drawn faded on the map, so they read as a rough area rather than a precise point.

Source: ISO/RTO interconnection queues (LBNL Queued Up) and LBNL interconnection cost studies. Withdrawal rates are empirical, per region; survival-model validation varies by region.