Driving through underground rocks
Steering a drill bit horizontally through rock is driving blind. You cannot see the layers, but you can measure the rock the bit is passing through, and a nearby vertical well gives you a map of the layers. Match the two and you know where you are. That is geosteering. This post is about a Kaggle competition on it: the one error a standard tracker cannot see, and how simulating the problem fixed it.
The setup
Two logs. Along the horizontal well, a gamma-ray log records how radioactive the rock is at each point the bit passes. From a nearby vertical well, a reference log records the same thing straight down through the layers, so it works as a lookup table: at depth in the layer stack, the rock reads . If the bit's reading matches the reference at some depth, the bit is probably at that depth.
The standard tool is a particle filter: keep a few hundred guesses of where the bit is and how steeply the layers are tilted (the dip), move every guess forward, score it against the newest reading, and keep the good guesses. With indexing steps along the well, the depth in the layer stack, the dip, and the gamma-ray reading, the model is three lines:
The first line says the reading should match the map at the current depth. The second says depth changes with the dip as the bit moves. The third lets the dip drift, but around a typical value . It works well and follows the layer boundaries closely. That last term matters later.
One number explains most of the error
Compared to the true path, the filter's error had a very simple shape: almost a straight line, small at the start of the well and growing with distance travelled, plus a smaller part with no trend.
Here is distance along the well, is one number per well, the gap between the filter's average dip and the true one, and is what remains: the shape error, which is what the filter actually spends its effort on. Remove and the error roughly halves.
The filter gets the dip wrong because of the term in its own model. The assumption about what a typical dip looks like quietly pulls every well toward that value. The pull is the same for every particle and every random seed, so running more of them does not help. Averaging seeds shrinks random error like , but is identical in every seed and survives the average untouched. It is a bias, not noise, and more compute buys nothing.
Why the filter cannot fix itself
The obvious repair is to let the filter grade itself: tilt the recovered path by a candidate , re-score it against the reference log, and keep the that fits best.
I tried it. The it picked had essentially no relationship to the true .
The reason is the useful part. The filter had already chosen to make this same misfit small, with the wrong dip baked in. So the misfit is flat around by construction, and small tilts all score about the same. On top of that, the leftover shape error is larger than the tilt you are looking for, so it swamps the signal. The signal is real, since running the same scan on the true path recovers immediately, but the filter's own fit hides it.
Put plainly: a model cannot audit itself with the objective it was fitted on. The check has to come from somewhere else.
Learn the error from simulation
So estimate from outside. Take a well with a known path , tilt it by a known , re-render the gamma-ray log the bit would have seen, and you have a training example with an exact label. Generate many of these and train a small network to recover the tilt from the rendering:
Then apply to real wells. This is simulation-based inference. Three things decided whether it worked, and none of them was the network . Each is one piece of the lines above.
- What the noise looks like (the ). My first simulator added random noise along the well. Real mismatch is different: the reference log is wrong about a particular layer, and wrong the same way every time the well crosses it. Until the simulator modelled noise that way, the network did worse than predicting nothing.
- What the network sees (the ). The raw gamma-ray sequence barely worked. A 2-D alignment image, showing how well the log matches the reference at every candidate depth and position, worked about ten times better. The information lives in that misfit surface, so hand it over rather than making the network find it.
- How wide you simulate (the ). With a narrow , the network never sees the badly tilted wells, which are the ones that matter most. Widening it is what made the estimates on the worst wells useful.
A separate calibration check also showed that the spread of was badly too narrow, even where its centre was fine. The point estimates were right. The confidence attached to them was not.
Bottom line
A well-built filter can be accurate and still carry a bias it cannot see, and neither more sampling nor self-checking will surface it. Simulation is a way out, but the work is in the simulator: realistic noise, the right representation, and a wide enough range of the quantity you want to learn. Get those right and a small network is enough.
References
-
Sequential Monte Carlo — Doucet, A., Godsill, S., & Andrieu, C. (2000). On Sequential Monte Carlo Sampling Methods for Bayesian Filtering. Statistics and Computing, 10(3), 197–208.
-
Simulation-based inference — Cranmer, K., Brehmer, J., & Louppe, G. (2020). The Frontier of Simulation-Based Inference. PNAS, 117(48), 30055–30062.
-
Simulation-based calibration — Talts, S., Betancourt, M., Simpson, D., Vehtari, A., & Gelman, A. (2018). Validating Bayesian Inference Algorithms with Simulation-Based Calibration. Preprint.
-
Amortised posterior estimation — Papamakarios, G., & Murray, I. (2016). Fast ε-free Inference of Simulation Models with Bayesian Conditional Density Estimation. NeurIPS.
-
Competition — ROGII: Wellbore Geology Prediction. Kaggle, 2026.
Other posts
- Can we really get alpha from market data?
The efficient market view, the micro alpha counter-argument, and why a weak signal only becomes a position once you know its uncertainty.
- What works for forecasting macro economic series with deep learning?
Korean output and investment nowcasting with seven deep models: what the data allows, which families worked, and why it depends on the target.
- Could multivariate time series have their own representations?
Why forecast embeddings are not factors, and how identifiable innovations with diagonal dynamics recover them without losing forecast quality.
- Can we make a more risk-aware portfolio agent from utility theory?
Epstein–Zin recursive utility inside actor–critic RL: the Bellman backup that changes, and what it did on Korean ETF splits.
- Classifying bird sounds in the field
Sound as a picture: what a spectrogram is, why the mel scale matches how we hear, and how a network finds a bird call in the image.
- Creating and Evaluating Synthetic Tabular Data
Sequential synthesis for credit bureau data, and three checks: pMSE distinguishability, confidence interval overlap, and attribute disclosure.