
Turning Measurements Into Worlds: The Physics and Visualization Behind TerraForge
NASA's exoplanet archive is one of the great open datasets in science, and it renders every planet as a point on a chart. TerraForge takes 5,574 of those rows and builds the world from the physics, while staying honest about where measurement ends and inference starts.
TerraForge now carries the NASA Exoplanet Archive. There are 6,359 planets in the current snapshot, and 5,574 of them have enough measured data to simulate, which means each one gets its own page: rendered, classified, scored for habitability, and sitting next to whatever a user built by hand that morning.

https://terraforge.chriswest.tech/exoplanets/ab-pic-b
Getting there meant confronting something I had been able to ignore while the app was a sandbox. When you let people invent planets, you can define the rules however you like. The moment you put a real planet on screen and label it, you are making a claim about an object twelve hundred light years away that somebody spent years measuring. You had better be careful about what you are claiming, and honest about what you are guessing.
That turned out to be the interesting engineering problem. Not the rendering. Knowing where the facts end and the guessing starts.
What NASA gives you, and what it doesn't
The archive is one of the great open data resources in science. It is operated by the NASA Exoplanet Science Institute at Caltech, it is free, it needs no API key, and it will hand you every confirmed planet, its host star, its orbit and its discovery circumstances through a public query interface. As I write this it lists 6,366 confirmed planets. It updates weekly. Six new ones went in the week I was finishing this piece.
What it does not do is show you the planet.
That is not a criticism. The archive is built for researchers, and its visual vocabulary is exactly right for that job: scatter plots of mass against orbital period, radius against mass, histograms of discoveries by year. Rigorous, legible, and completely abstract. A planet is a point.
NASA does visualize this, and does it well. Eyes on Exoplanets flies you across the galaxy, renders every known system in 3D with its habitable zone drawn in, and lets you hold any of them up against our own solar system. It runs on the same archive I am using. I had a version of this section that claimed nobody was doing this, and it was wrong. That tool covers a lot of the ground TerraForge covers.
What it shows you for each planet, in NASA's own words, is an artist's concept. Someone read the parameters and painted an interpretation of what the world might look like. That phrase turns up all through the documentation, and it is doing deliberate work. An artist's concept announces itself as imagination.
Which is arguably the more honest choice, and I want to sit with that before describing mine as an improvement. TerraForge generates its planets from the numbers. That sounds more rigorous, and in one narrow sense it is: the same inputs always produce the same world, and when a measurement gets revised the picture changes with it. But an image that fell out of a model can carry more authority than it has earned. Nobody mistakes a painting for a photograph. People absolutely mistake a simulation for a measurement.

TerraForge's Exoplanet Gallery
So the real difference is not that they illustrate and I compute. It is that Eyes on Exoplanets is a viewer and TerraForge is a sandbox. Theirs shows you what is known, and shows it beautifully. Mine lets you take a real catalogued planet, change its composition or move it closer to its star, and watch the classification come apart in your hands. You can build a world that does not exist and ask which real one it most resembles.
Scoring a world

Kepler 442b : https://terraforge.chriswest.tech/exoplanets/kepler-442-b
The layer neither of them has is judgment. The archive hands you parameters. Eyes draws a habitable zone, which is a band of orbital distance where liquid water could exist. TerraForge scores every world on seven weighted factors: temperature, atmosphere, water, magnetic field, geology, chemistry and rotation. Those roll into a single number from 0 to 100 and a band running from Uninhabitable through Marginal up to Highly Habitable, and every factor shows its own reasoning. Why a world lost points for having no magnetic field to hold an atmosphere down. Why a chlorine-heavy mix wrecks the chemistry score, chlorine being efficient at taking organic molecules apart. Why a rotation measured in months produces temperature swings nothing survives.
Each of those factors is its own small piece of physics rather than a slider I set by feel. Atmosphere retention is a tug of war between escape velocity and how fast the gas molecules are moving: a hot, low-mass world loses its hydrogen because the molecules outrun the planet's grip, which is why small worlds close to their star end up bare rock. Magnetic field strength needs a conducting liquid core and enough rotation to stir it, so it falls out of iron content, mass and day length acting together. Geology runs on heat, which comes from radioactive elements in the mix plus whatever the planet kept from its formation, and a large planet stays hot for far longer than a small one. Water depends on both having hydrogen and oxygen present and on the surface sitting in the narrow temperature band where the result is liquid rather than ice or steam.
Temperature is the one that does the most work. It starts from the equilibrium value, then takes a greenhouse correction weighted by what is actually in the atmosphere, because gases differ enormously in how much warming each one buys per unit present. This is not an exotic idea: Earth runs roughly 33 K warmer than its bare equilibrium temperature for precisely that reason, and Venus is the cautionary tale at the other end. Get that correction wrong and a temperate world reads as frozen, or a lava planet reads as habitable.
I am going to stay vague about the specific weights, thresholds and curve shapes, since those are what I have spent the most time tuning and they are the part worth keeping. But the shape of it is not a secret. Every factor is a physical mechanism fed by inputs the app already has, and the score is simply what falls out when you make each mechanism produce a number and then decide how much each one should count.
That scoring is the thing people actually engage with. It turns a row of measurements into a verdict you can argue with, and arguing with it is the point: move the planet, change the mix, watch the number move and read why.
I should also be straight about where it sits on the confidence scale, because it is the softest thing in the entire app. Composition inference is at least anchored to a measured density. A habitability score is an opinion with arithmetic attached. I chose the seven factors. I chose their weights. I chose the thresholds between the bands. There is no validated habitability metric to check any of it against, because we have exactly one confirmed example of an inhabited planet and it is the one we are standing on. The score is genuinely useful for ranking worlds against each other and for making the physics legible to someone who is not going to read a mass-radius curve. It is not a prediction that anything lives there.
That is a different question from what is actually out there, and it needs different machinery. It is also the reason the NASA data earned a place in a planet-building toy: not because the archive needed prettier pictures, but because once real worlds sit next to invented ones, you can finally ask how far a real planet is from the thing you imagined.
Two numbers and a guess
Here is the uncomfortable fact at the center of the whole thing: nobody can measure what an exoplanet is made of.
There is no catalog column for element abundances, because there is no instrument that returns them. What we actually have, for the planets we know best, is two measurements. Transit photometry gives you a radius, from how much starlight the planet blocks. Radial velocity or transit timing gives you a mass, from how much the planet tugs its star around. That is close to the whole inventory.
From those two you get bulk density, and density is genuinely informative. A planet with Earth's mass packed into half Earth's volume is made of heavy things. A planet with eight Earth masses spread across three Earth radii is holding a lot of something light. That much is real inference from real measurement.
What density cannot tell you is which light thing. A given density is satisfied by a lot of different mixtures, and picking one is where the modeling starts and the certainty stops.
Three curves and the space between them
The model in TerraForge is built from three reference mass-radius relationships, each describing a planet made purely of one thing. A rocky world, using a Zeng-style fit for an Earth-like iron and silicate mix, follows roughly R = M^0.27, calibrated so that Earth lands exactly at one Earth radius. A water and ice world sits higher. A hydrogen and helium envelope sits much higher, with an exponent close to flat, because hydrogen is so compressible that piling on mass barely widens the planet.
Everything real lives between those curves. So the app mixes them by volume additivity and binary-searches for the fraction that reproduces the observed radius. Give it a mass and a radius, and it hands back something like "about a third water by mass."
I calibrated the constants against planets where the literature has an independent answer. TOI-1452 b is the useful one: a discovery paper puts it near 30 percent water by mass, and the model has to land in that neighborhood. GJ 1214 b and K2-18 b have to come out water-rich. A set of well-characterized rocky planets, including several TRAPPIST-1 worlds and 55 Cancri e, have to come out rocky. When a constant change broke one of those, the constant was wrong.
The part where the field disagrees with itself
Then I did the thing I should have done earlier, which was check my classifications against the actual literature rather than against my own intuition.
The method held up. Luque and Pallé (2022) separate exoplanet populations using density normalized against an Earth-composition curve, which is essentially what TerraForge does. Their approach is the field's approach, and mine is a version of it.
The results were less comfortable. By their metric, TOI-1452 b, one of my anchor planets, sits on the rocky side of the density gap rather than the water-rich side. And the water-world population itself is contested: later work, including Dainese and Albrecht (2025), roughly doubled the sample and found two populations where Luque and Pallé found three. Grading my own curves against the sources, the rocky exponent is right but my coefficient runs a few percent high, my single water curve stands in for what is really a temperature-dependent family of curves, and my gas curve puts giants at roughly half the radius the standard models give.
I wrote a line in my own notes while working through this that I keep returning to: the model picks one composition where the field says the data cannot. That is an honesty problem before it is an accuracy problem.
So the fix was not better constants. It was better labels. A planet inferred to be water-rich now says so in language that admits it is inferred. Rows whose radius comes from an archive model rather than a transit say "estimated" instead of "measured," because those two words are doing real work. And a few approaches got investigated and thrown out along the way, including one that tuned the curve constants until the classifications came out nicer, which is a polite way of describing throwing away your validation set.
The bug report that was wrong
My favorite episode in all of this was a bug that turned out not to be one.
TerraForge computes equilibrium temperature with the standard relation, scaled by stellar luminosity and orbital distance. The reported bug was that the formula ignores albedo, the fraction of light a planet reflects. Add a realistic albedo term, the report said, and the numbers will be right.
It would have been an easy fix. It was also wrong. The archive's own equilibrium temperature column is defined as the temperature of the planet modeled as a black body heated only by its host star, and a black body by definition reflects nothing. Adding an albedo term would have moved the model away from the reference it is supposed to match. I checked what the change would have done across the catalog and it would have reclassified hundreds of worlds in the wrong direction.
What made this checkable was having something to check against. Comparing TerraForge's temperature to the archive's across more than five thousand rows, the ratio has a median of about 1.03 and a spread from roughly 0.97 to 1.18. That spread is the honest answer: no single albedo value could correct it, because the disagreement is not a constant offset. It is the model and the archive making different assumptions about different planets.
I kept the zero-albedo formula and wrote down why.
How this actually gets built
I should say plainly how a front-end engineer with an art degree ends up shipping mass-radius curves and orbital mechanics. I build this with AI, heavily, and that includes the physics and the math.
I am not an astrophysicist. I can read a paper and follow the argument, but I could not derive a volume-additivity mixing law from a standing start, and I could not tell you offhand which exponent belongs on a water-ice curve. Working with AI got me from "I want composition inferred from density" to running code far faster than I could have managed alone. It is the reason a side project has an actual model in it rather than a lookup table of nice-looking colors.
Which is exactly why everything above is built the way it is.
AI is very good at producing something that looks like physics. It will hand you a formula with confident commentary attached, and the formula will be plausible, and plausible is not the same as correct. The albedo episode is the perfect illustration. That suggestion was reasonable, well argued, and wrong, and the only reason it got caught is that there was an authoritative definition to check it against and thousands of rows to test it on.
So the guardrails are the real work, and they are the part I own. Calibration planets with published answers that the model has to reproduce. A test suite that fails loudly when a constant change breaks a known world. A comparison against NASA's own values across the entire catalog. And eventually the slow, unglamorous business of reading the actual papers and grading my own curves against them, which is how I found out my water curve was standing in for a family of curves and my giants were half the radius they should be.
AI wrote a lot of this math. Verification is what makes it defensible, and no model gets to skip that step because it sounded convincing on the way in.
Small honesties
Most of the integrity work is unglamorous and lives in details nobody will notice, which is roughly the definition of doing it properly.
Star spectral class is derived from measured stellar temperature rather than the archive's spectral type column, because temperature is populated for about 95 percent of rows and spectral type for about 37. When you ask the app to find the real planet most similar to the one you built, it compares across mass, radius, temperature, insolation and host star in log space, and if nothing clears the similarity threshold it tells you your planet is one of a kind rather than forcing a bad match.
The new spectroscopy view shows real absorption lines at their real wavelengths, so hydrogen alpha sits at 656.3 nanometers where it belongs. Real exoplanets display their actual measured distance and discovery method. Planets you invented get an explanation of how parallax works and no distance number at all, because inventing one would be the exact failure the rest of this was built to avoid.
The catalog refresh is a scheduled job that opens a pull request rather than pushing to main, and the fetch script refuses to write if the row count drops by more than five percent. A renamed column upstream should fail loudly, not quietly delete several thousand pages.
Seeing it
None of this matters if the result still looks like a spreadsheet. Closing the distance between a row of numbers and a place was the entire point, so the visual side got the same attention as the physics.
Surface view drops you onto the planet, rendered entirely from that world's own numbers. Terrain, oceans, sea ice, lava, vegetation and weather all derive from the composition and temperature you configured or the archive reported. There is a hard rule in that part of the codebase that nothing calls the random number generator directly, because the same planet has to look the same every time you visit it. Randomness is fine. Non-determinism is not.
The galaxy is the piece I am most attached to. Every public star system has a permanent home in a rendered spiral galaxy built from a quarter of a million particles, and you can fly to any of them and arrive at a live orrery of that system. A system's location is seeded from its identifier and written once, so a world someone built in July is exactly where they left it.
The homepage got rebuilt around that community rather than around me. Real user-made worlds in the hero, a running count of everything forged so far, and freshness fixes so that a planet published a minute ago actually appears when you go back.
What the model is
I want to end where the engineering ended up, which is a more modest place than it started.
TerraForge does not tell you what an exoplanet is made of. It tells you what composition is most consistent with two measured numbers under a specific set of assumptions, and it is increasingly careful to say so. The science is real: the curves come from published models, the data comes from the archive, the calibration comes from planets where somebody did the hard work of measuring.
But the output is a hypothesis with a number attached, not a fact. The difference between those two things is most of what I learned building this, and it is the part I am proudest of getting into the interface rather than just into my notes.
Comments
Loading comments...
