An open benchmark

Can AI read a birth chart?

Large language models describe a birth chart almost perfectly. Whether they can actually read one, and tell you something true about the person it belongs to, is a completely different question. That is the one we measure.

Vedic astrology, Parashari method, whole sign houses. Every model gets the same chart in the same format.

Results are being rerun. Here is what we know so far.

We published a first set of results and then withdrew them, because the way we scored the models measured the wrong thing. That page is gone. This one describes what the benchmark asks, what survived, and what is coming.

Facts they get right about the chart 95.5–100% This is the one number that survived unchanged. Our own software checks every astrological statement in a reading against the real chart, with no AI involved. Models describe the chart almost perfectly.
Whether the reading is true of the person Being measured The harder half, and the reason this benchmark exists. Reading a chart correctly and saying something true about a life are not the same skill, and only the first one is easy to check.
A ranking of models Withdrawn We published one and then withdrew it. It ranked named products using a score that measured the wrong task, so it told you nothing useful about any of them.

A real question, asked the way people actually ask it

Nobody sits down with an astrologer and asks for a percentage. They ask whether they will marry, whether the move abroad will happen, when the work will finally turn. So that is what we give the model: a chart, a date, and a question in a person's own words. It writes a reading in prose, the way a practitioner would.

The chart

The model gets the chart and a consultation date, and nothing else.

It never sees a name, a birth year or a birthplace. If it could recognise the person, it would score well without reading anything.

The question

Asked before the event, not after.

We date the consultation ahead of what actually happened, so the model is predicting rather than describing a finished life. Asking someone at 30 whether marriage is coming is a real question. Asking a completed biography when the marriage was is not.

The verdict

Right, partly right, wrong, or not answered.

That last one carries a lot of weight. A reading saying there may be difficulty but matters will resolve has committed to nothing and scores nothing. Without it, vague readings quietly collect half marks and every model looks capable.

Our first attempt measured the wrong thing

We asked models for percentages and graded them with statistics. Nobody consults an astrologer that way, and it turned out not to be a sound way to test a language model either.

We had published results from it. When we understood what the scoring was actually measuring, we took them down rather than leave them up with a note attached. They are not coming back, and they should not be quoted.

We are saying so here rather than quietly deleting the page. A benchmark that cannot tell you when it got something wrong is not worth reading.

Questions people ask about AI and birth charts

Can AI read a birth chart?

It reads the chart itself very accurately. Across every LLM we have tested, 95.5% to 100% of the factual statements a model made about a chart were true of that chart: which planet sits where, which sign rules what, which period was running. Whether it can turn that into a correct statement about the person is the open question, and it is what we are rebuilding the test to answer.

Which LLM is best at astrology?

We do not have an answer we are willing to stand behind. Our first attempt produced a ranking, but it scored models in a way we have since decided was wrong, so we withdrew it rather than leave a league table up that measured the wrong thing. The rerun covers Claude, GPT, Gemini, Grok, DeepSeek, Kimi, Qwen, GLM and Muse models.

Does this prove or disprove astrology?

Neither, and we are careful never to claim it does. This is a test of AI models, not of astrology. No human astrologer takes part, and no benchmark of language models could settle whether the underlying system works. We measure something much narrower: whether a birth chart gives an LLM anything it can actually use.

How do you check whether an AI reading was right?

Two separate checks. First, our own software verifies every astrological statement in the reading against the actual chart, with no AI involved. Second, a judge who holds the real outcome marks the reading right, partly right, wrong, or not answered. That last category matters: a reading that says difficulties may arise but will resolve with effort has committed to nothing, and it earns nothing.

Why did you retire your first method?

It asked models for percentages and graded them with statistics. That is not how people consult an astrologer, and it is not a sound way to test a language model. Real people ask a question and get an answer, so that is what we now test: a chart, a date, a question in someone\u2019s own words, and a reading we can check against what actually happened.

What does this mean if I am building an AI agent?

One finding has held through every version of this test and is worth borrowing. These models handle the part that looks hard, reading a dense and unfamiliar data format, almost perfectly. Checking that your agent parsed its inputs correctly will therefore tell you almost nothing, because that check passes every time. The failures live in the conclusions drawn afterwards, and the only thing that catches them is scoring those conclusions against something real.

When will results be published?

When they are worth publishing. The current run is deliberately small while we validate the judging, and a benchmark that publishes early is how a leaderboard ends up ranking noise. We would rather be late than post a table we have to withdraw, which is exactly what happened the first time.

Know someone who would argue with this?

Send it to them.

X LinkedIn