Executive summary
Every US onshore operator, investor, and service company relies on the same underlying facts: which wells exist, who operates them, how they were completed, and how much they have produced. Those facts are public. Six state regulatory agencies publish them. And yet, turning that public record into an answer you can put in front of an investment committee still takes hours of downloading, cleaning, joining, and normalizing — per question.
WellStrata is a data platform that closes that gap. It ingests raw regulatory filings from six onshore states, scores every record for quality, extracts completion parameters that never appear in bulk downloads, and exposes the result through an AI query interface and a set of purpose-built analytical tools. This paper describes the data foundation, the quality-control methodology, and the analytical layer — and, as importantly, where the empirical data ends and engineering judgment must begin.
Coverage at a glance: 1,200,000+ wells across Texas, Oklahoma, North Dakota, New Mexico, Colorado, and Wyoming, refreshed monthly, with a 0–100 quality score on every record.
The problem with public data
Public oil and gas data is not hard to find. It is hard to use.
The traditional path to a single answer — say, "Which Wolfcamp B wells in Reeves County had the highest six-month cumulative oil after 2020?" — runs through a state portal, a bulk CSV export, a spreadsheet, a filter, a join against a completion table, a pivot on operator name, and a cleanup pass to reconcile the fact that one operator appears under six different name variants. Thirty minutes later you have something resembling an answer. Then you realize the state's bulk export lags real production by 45 days.
Four structural problems recur across every state:
- Fragmentation. Each agency — TX RRC, OK OCC, ND DMR, NM OCD, CO ECMC, WY WOGCC — publishes on its own schedule, in its own schema, with its own identifiers. Nothing joins cleanly across state lines.
- Latency. Bulk exports lag current production by weeks. Decisions made against stale data carry silent risk.
- Inconsistency. Operator names, formation aliases, and units are not standardized. "Wolfcamp," "WOLFCAMP A," and "Wfmp A" may all refer to distinct — or identical — targets depending on the filer.
- Missing depth. The parameters that actually drive well performance — proppant loading, fluid type, stage count, perforation intervals — live inside scanned completion PDFs, not in any bulk table.
The result is predictable: engineers skip analysis that would take too long relative to the value they expect from it. The most useful questions are often the ones that get abandoned.
The data foundation
WellStrata's ingestion pipeline pulls from each state's regulatory system on a monthly cadence and consolidates six independent schemas into one normalized model.
| State | Agency | Key formations |
|---|---|---|
| Texas (TX) | Railroad Commission (RRC) | Wolfcamp, Bone Spring, Spraberry, Delaware |
| Oklahoma (OK) | Corporation Commission (OCC) | SCOOP, STACK, Granite Wash |
| North Dakota (ND) | Dept. of Mineral Resources (DMR) | Bakken, Three Forks |
| New Mexico (NM) | Oil Conservation Division (OCD) | Delaware, Wolfcamp, Bone Spring |
| Colorado (CO) | Energy & Carbon Management Comm. (ECMC) | Niobrara, Codell, Wattenberg |
| Wyoming (WY) | Oil & Gas Conservation Comm. (WOGCC) | Powder River, Green River, Turner |
The consolidated model covers well headers and locations, monthly production and injection histories, completion filings, and hydraulic-fracturing chemical disclosures (FracFocus). Every record carries data-freshness metadata so a downstream consumer always knows the vintage of what they are looking at — the raw source and the QC layer are both preserved, so nothing is lost in normalization.
Quality control: a score on every record
The distinguishing feature of the platform is not that it aggregates public data — it is that it tells you how much to trust each data point. Every record is assigned a composite quality score from 0 to 100, bucketed into four bands:
| Band | Score | Interpretation |
|---|---|---|
| Gold | 90–100 | Complete, internally consistent, spatially validated |
| Silver | 70–89 | Reliable, minor gaps or unreconciled fields |
| Bronze | 50–69 | Usable with caution; notable gaps |
| Flagged | < 50 | Anomalies detected; verify before use |
The score is produced by several independent checks:
- Canonical operator names. Operator name variants are resolved to a single canonical entity. Without this, any per-operator ranking is quietly wrong — an operator split across six spellings appears as six small operators instead of one large one.
- Coordinate validation. Surface and bottom-hole locations are checked for plausibility against the reported county, state, and formation footprint. Wells that plot in the wrong county — or in the ocean — are flagged rather than silently mapped.
- Anomaly detection. Production and completion values are screened for physically implausible readings and step changes that indicate reporting errors rather than real events.
- Completeness. Records missing fields required for downstream analysis (lateral length, first-production date, formation) score lower and are labeled accordingly.
Because the score is a first-class field, it is queryable. An engineer can ask for "Bone Spring wells with a QC score above 80 in Eddy County" and exclude the noise before analysis begins — rather than discovering data problems halfway through building a type curve.
The most valuable thing a data platform can tell you is not just what the number is, but how much to trust it.
Document intelligence
Some of the most decision-relevant parameters are not in any bulk download — they are locked inside regulatory completion reports filed as PDFs. WellStrata applies document intelligence to extract structured completion data from Texas W-2 and New Mexico C-102 forms, recovering:
- Stage count
- Proppant volume
- Frac fluid type
- Perforation intervals
This turns a scanned form into a queryable field, surfacing completion detail that competitors relying on bulk exports simply do not have.
The analytical layer
Clean, scored, normalized data is the foundation. The value shows up in the tools built on top of it. WellStrata's modules are designed to hand off to one another — a question answered in one becomes the starting filter for the next.
AI Query: the first layer of understanding
The AI Query interface is a chat-style input. You type a question in plain English and receive a structured answer — a natural-language summary, a ranked table, or a production chart, depending on what you asked. A domain-aware language model interprets the question (it knows what IP30 means, that "Wolfcamp" is a formation, that "cumulative oil" differs from "monthly production"), translates it into a structured query against the database, and returns results in the most sensible format.
It handles the questions engineers and land analysts actually ask: formation comparisons, operator rankings, completion benchmarks, QC-filtered result sets, and state-level rollups with production thresholds. A well-formed query typically consumes 500–2,000 tokens; usage is tracked transparently in the interface.
The point is not to replace careful analysis. It is to drop the cost of the first question so low that questions which used to be abandoned get asked.
Wells Explorer: geography without county-line noise
The Wells Explorer renders every well in the database as formation-colored dots on a fast interactive map. Its most useful capability is polygon selection: draw any shape — an acreage block, a prospect radius, a segment that straddles a county line — and every well inside the geometry is selected, loaded into a table, and exportable. Acreage does not respect county boundaries, and formation-plus-county filtering gives too much noise for real screening work.
Clicking any well opens a detail panel with its full production history and its FracFocus disclosure (water volume, TVD, and the complete ingredient list, with CAS numbers hyperlinked to PubChem). Filter states are deep-linkable as URLs and savable as named screens.
Operator Benchmarking
The Benchmarking module ranks operators in a basin across four metrics — IP30, cumulative oil, well count, and average lateral length — with a minimum-well-count filter to separate genuine performance from small-sample luck. IP30 values are color-coded by quartile; the table is sortable by any column; and any operator can be drilled into for their production trend, top wells, and formation mix.
Crucially, average lateral length sits alongside IP30 in every view — because an operator running 12,000-ft laterals will show higher raw IP30 than one running 8,000-ft laterals in the same county, even if the shorter-lateral operator is doing the better completion per foot. The tool surfaces the context; the judgment stays with the engineer.
Formations Analytics
Basin-level numbers blend distinct plays together. The Wolfcamp A and Wolfcamp B are not the same rock; the Niobrara and Codell within the DJ Basin have different profiles and different trajectories. Formations Analytics works at the formation level, loading four charts at once: monthly production history, wells drilled per year, top operators by cumulative oil, and — the most analytically valuable — average IP30 by year with a linear regression trend line.
That regression cuts through vintage-to-vintage noise to show direction. A formation where average IP30 climbed from 400 bbl/d in 2015 to 750 bbl/d in 2024 is a fundamentally different investment thesis than one that peaked in 2019 and has declined since.
Chemicals Explorer
FracFocus has collected hydraulic-fracturing chemical disclosures since 2011 — millions of job disclosures, publicly available, and almost never analyzed systematically because the raw structure resists cross-well analysis. The Chemicals Explorer processes it into three views: Operator Comparison (water per job, ingredient profiles), Ingredient Explorer (who uses a given chemical, and the year-over-year adoption trend), and Water Intensity (basin-level water-per-job trends). Every CAS number links to its PubChem entry. The data supports completion benchmarking, vendor research, ESG and water-management analysis, and technical due diligence — all from records that are public by law.
Worked methodology: a type curve you can defend
Type curves feed EUR estimates, which feed reserve reports, which feed acquisition pricing and capital allocation. Given the stakes, it is remarkable how often they are built under time pressure from incomplete peer sets. WellStrata's Type Curve Builder is designed to make the rigorous version fast enough to be the default.
A defensible type curve depends entirely on peer selection. The builder filters on the variables that matter:
- Formation — the most important filter. Wolfcamp A and B have different pressure regimes and completion responses; mixing them represents neither.
- Geography — county or sub-county. Formation name alone is not enough; the Wolfcamp in Midland County is different rock than in Reeves County.
- Vintage — completion-year range. Frac designs improved materially from 2015 to 2022.
- Lateral length — filter to a comparable range, or normalize rates per 1,000 ft of lateral.
- Minimum production history — a well with six months of history cannot tell you what month 24 looks like.
The builder aligns every matching well to months-on-production (MOP 1 = first full calendar month), then calculates P10, P50, and P90 rates at every month. It plots the P50 as a bold median line inside a shaded P10–P90 band, reports full-life EUR at each percentile — fitting a modified-hyperbolic decline out to a terminal exponential rather than summing only the visible months — and shows the well count contributing to each MOP. A multi-curve overlay turns it into a range-of-outcomes analysis: pin a base case, change one filter, and compare 2018 vs. 2022 vintage, or long vs. short laterals, or an operator against the formation P50.
What the curve cannot tell you
An honest whitepaper names its limits. A type curve represents the historical performance of the peer set. It does not account for:
- Future commodity prices or operating-cost changes
- Planned completion-design changes with no analog in the peer set
- Subsurface heterogeneity not captured by the formation-county filter
- Parent-child interference at a specific location
These are engineering-judgment adjustments that belong on top of the empirical curve, not inside it. The platform gives you a defensible base and the exact filter criteria that generated it — so any reviewer can evaluate the peer selection before debating the EUR.
Where the data ends
WellStrata is deliberately not a reservoir simulator, a GIS system, or an economics package. The Wells Explorer has no projection controls or raster overlays; the AI Query interface does not do reservoir modeling; the type curve does not price a deal. What the platform provides is the trustworthy, queryable, decision-ready base layer — every onshore well, scored for quality, with the completion and production context a working engineer needs to answer the question at hand without leaving the browser.
For teams that want machine-learning forecasts, refrac screening, and IOR/EOR ranking on top of this same data, BIROVA's Resery models are consumable inside WellStrata premium plans via API. For GPU-accelerated reservoir simulation, see Permeon.
Conclusion
The public record already contains the facts every onshore decision depends on. The bottleneck has never been availability — it has been trust, latency, and the hours of manual work between a raw export and a defensible answer. WellStrata removes that bottleneck: six states, 1.2M+ wells, a quality score on every record, monthly refreshes, and an analytical layer that hands off cleanly from a plain-English question to a reserve-review-ready type curve.
The most honest analysis shows its work. WellStrata is built so that it can.
WellStrata is a product of BIROVA. All capabilities described are available via API and web UI; the 14-day trial includes full access across all six states with no credit card required. Figures reflect current coverage and are refreshed monthly from state regulatory agencies.