brihat.ai
← Writing

Breaking the TTV Degeneracy: I Said the Fix Was a Second Observable. Here Is the Test.

2026-06-18

In the last post I built an amortized, calibrated posterior estimator for exoplanet masses from transit-timing variations, and reached a conclusion that was as much a limitation as a result. The network is honest, even across the 2:1 resonance, but near resonance its posteriors get wide. The mass-eccentricity degeneracy means a given timing signal is consistent with a heavy planet on a near-circular orbit or a lighter one on a more eccentric orbit, and an independent reference posterior confirmed that this width is physical: the transit times genuinely do not contain enough information to separate the two. The implication was a prediction. If the limit is the data, then a better estimator will not help, but a second, complementary observable should.

This post tests that prediction. The result is clean enough that I want to show it on its own: adding one observable collapses the degeneracy.

The observable: transit durations

So far the network saw only transit times, summarized as the O-C residuals (each transit's deviation from a perfectly periodic clock). A transit also has a duration, the time the planet takes to cross the stellar disk. In the edge-on, coplanar geometry of the model, the planet crosses a chord of length 2·R_star at its sky-plane speed at mid-transit, so

duration  =  2 · R_star / v_sky

The useful part is what sets v_sky. A planet transits at the same orbital longitude every time, so its transit speed barely changes from transit to transit. What changes the speed is the planet's eccentricity: an eccentric orbit moves faster or slower than a circular one at a fixed longitude, which lengthens or shortens the duration by up to a few tens of minutes for the eccentricities here. Crucially, that speed depends on the planet's own orbit, not on the mass of the perturber. So the duration carries a near-direct readout of the eccentricity, almost independent of mass.

I checked this in the simulator before training anything. Referenced to the circular-orbit duration, the per-system duration anomaly correlates with the eccentricity-vector component h = e·cos ϖ at essentially −1.00, and the signal is 30-ish times the assumed timing noise. Timing constrains roughly mass × f(eccentricity); duration constrains the eccentricity directly. Together they should pull mass and eccentricity apart.

The experiment

To isolate the effect of the observable from everything else, I trained two identical mixture-density networks on the same systems and the same noise draws, changing only the feature:

  • Arm A, timing only: the 60 O-C residuals (40 inner transits, 20 outer).
  • Arm B, timing + durations: those 60, plus 60 duration anomalies.

Same architecture, same optimizer, same training budget. Then I measured, on a held-out test set, how tight each posterior is relative to the prior (a ratio near 1 means the data taught the network nothing; near 0 means it pinned the parameter down) and whether calibration survived.

The result

Adding durations sharpened every parameter, and the masses most of all:

parameter timing only timing + durations tightened by
m₁ 0.46 0.11 76%
m₂ 0.43 0.08 81%
h₁ 0.23 0.01 94%
k₁ 0.21 0.07 66%
h₂ 0.30 0.01 95%
k₂ 0.29 0.12 58%

(Values are posterior width divided by prior width; smaller is sharper.) The best validation likelihood improved by about 5.9 nats, the information-theoretic version of the same statement. And 90% coverage stayed near nominal in both arms (0.85 to 0.92 across parameters), so the sharpening is genuine information, not the model becoming overconfident.

The mechanism is exactly the predicted one. The h components, which the duration measures almost directly, tighten the most (94 to 95%). Pinning the eccentricity collapses the degeneracy, so the timing amplitude finally resolves into a mass, and m₁ and m₂ tighten by about 80%. The k components, the eccentricity direction the duration does not probe, tighten the least. The pattern is the fingerprint of the physics, not a generic "more inputs help."

The degeneracy, before and after

Here is the same single system, inferred both ways. The truth is the star.

Posterior samples for one system. In violet, timing only: the broad, tilted mass-eccentricity ridge. In cyan, timing plus durations: the same system and network, with one extra observable, collapsed to a tight blob on the truth. For this system the outer-planet mass uncertainty fell from about ±7.6 to ±0.5 Earth masses.

That violet smear is the degeneracy that defined the problem in the first post. The cyan blob is what happens when you give the network a measurement that constrains eccentricity on its own. The ridge does not get modeled better; it stops existing, because the data now distinguishes the points along it.

Near the resonance: a partial fix

The experiment above is at a single period ratio, comfortably away from the separatrix. But the whole point of the project was the regime across the 2:1 resonance, where in the first post the posteriors collapsed toward the prior. So I ran the same A/B there: one model spanning ratio 1.90 to 2.20 with larger eccentricities (e_max 0.15) and the measured ratio supplied as a conditioning input, averaged over three seeds. This is the setting where timing-only informativeness fell apart.

Durations help, and the way they fall short is the interesting part. Two things clearly recover. The lost information comes back: best validation NLL improves from +3.6 to −1.4, about five nats, on every seed. And the eccentricity h-components are pinned almost perfectly, width ratios dropping from 0.45 and 0.77 to 0.03. Calibration holds throughout. But the mass posterior tightens only about 22%, far less than the ~80% in the clean fixed-ratio case, and the improvement is weakest right at the separatrix.

The reason is a blind spot I should have seen coming. A duration measures the planet's speed at the transit longitude, which is the eccentricity-vector component h = e·cos ϖ. It says almost nothing about the perpendicular component k. Away from resonance, pinning h is enough to resolve the mass. But near resonance, with larger eccentricities, the mass couples to the full eccentricity vector, and with k still unconstrained the degeneracy is only half broken. The h-components collapse to a ratio of 0.03; the k-components barely move (k1 by 4%, k2 not reliably at all). So the mass relief is real but partial.

This sharpens the thesis rather than overturning it. More observables is still the lever, and durations confirm it by restoring information and eccentricity precision even across the separatrix. But "more observables" is not "any one observable." Durations break the h half of the degeneracy; fully closing the near-resonant mass would need something that constrains k as well, radial velocities, or a longer baseline with more transits. The next observable has a job description now.

What this does and does not settle

It confirms the central claim of the project: near resonance the bottleneck was observational, not algorithmic. The wide posteriors were the honest consequence of missing information, and supplying that information, not a better density estimator, is what sharpens them. It also shows the path forward is to enrich the forward model, the duration anomaly is computed from geometry the integrator already produces, so this was a change to what we measure, not to how we infer.

One caveat I can now retire. The headline experiment gave the durations the same 0.5-minute precision as the mid-transit times, which is optimistic, since real transit-duration uncertainties are coarser. So I ran a robustness sweep: keep the timing precision fixed, but degrade the duration noise to 2, 5, and 10 minutes, across three seeds. The effect survives. Even at 10-minute duration precision, twenty times coarser than the timing, the outer-planet mass posterior is still about half the width of the timing-only baseline (ratio 0.22 versus 0.43), and 90% coverage holds at 0.86 to 0.88 throughout. At a more typical few-minute precision it is nearly as sharp as the optimistic case. The reason it degrades so gently is that the eccentricity signal lives in the mean duration anomaly, so averaging over the 40 inner and 20 outer transits suppresses per-transit noise by about the square root of the count. The observable is robust because it is an averaged quantity.

The remaining caveats are narrower now. The before-and-after figure is a single system and seed (the sweeps are three). And the near-resonance test names the real open problem: an observable that constrains the k eccentricity component, so the mass degeneracy closes fully across the separatrix and not just in the circulating wings. That, plus swapping the mixture-density network for a normalizing flow to stabilize training in this regime, is where this goes next.

Update, June 2026. The observable that closes the k half turned out to be the obvious one: radial velocity. The star's reflex curve has an amplitude set by the planet's mass and an eccentric shape that fixes the full eccentricity vector, so it pins k where the duration cannot. Added to the same across-resonance test, the near-resonant outer-planet mass tightening jumps from the 22% above to about 80%, back to the fixed-ratio level, and the k-components finally move (around 50%) while the h-components stay pinned and calibration holds. None of this is a surprise. Radial velocity measuring planet masses is the oldest trick in the field, which is exactly why it stays a footnote here rather than its own post. The one wrinkle worth a sentence is observational: because the mass lives in the RV amplitude but k lives in the curve's shape, three RV epochs already buy half the mass constraint, while pinning k needs a denser campaign of around thirty. The scripts (observables_rv.py, robustness_rv.py, cadence_rv.py, posterior_overlay_rv.py) are in the same repo.

The code, the simulator change, and the scripts behind these numbers (observables_experiment.py, posterior_overlay.py, robustness_durations.py, and observables_cross_resonance.py) are at github.com/brihat9135/ttv-experiment.