How it works
A results press release goes in. A decision model answers a fixed set of questions about it. The answers come out as probabilities, which is what makes one release comparable with the next.
From filing to reading
- The text. When a US company reports results it files a Form 8-K with the SEC under Item 2.02, with the press release attached. That press release is taken from EDGAR.
- Cleaning. Financial tables and legal boilerplate are removed: the narrative is what gets read. Company names, tickers and dates are masked, so the model judges what the text says and not which company says it.
- The questions. The model is asked 10 fixed questions, the same for every company and every quarter. They are listed below, word for word.
- The answers. Each answer is a probability: how likely it is that guidance was raised, that margins are described as under pressure, that tariffs are said to affect the business.
The model
The reader is pplx-decider-v1-27b, Perplexity's Decider. It is a decision model, a different kind of tool from a chat model. A chat model is given a prompt and writes text. A decision model is given a text and a typed question, with the possible answers spelled out, and returns a probability for each of them. It cannot ramble, change the subject or invent a sentence, because it has no way to produce one.
How it is built
It starts from Qwen3.8-27B, an open general-purpose language model, and changes it in three ways.
- The output layer is replaced. A language model ends in a layer that picks the next word out of a vocabulary of some 250,000 tokens. Here that layer is taken out and a small decision readout put in its place: a single linear layer with 255 outputs, one per possible answer option. It starts as a copy of the rows the original layer used for the option letters, so the model begins from what it already knew about answering multiple-choice questions.
- The whole network is fine-tuned on decisions. The backbone and the readout were trained together, with supervised learning, on a curated set of 73,000 examples of a text, a question and the right answer. The released weights are the checkpoint taken after about 51,000 of them.
- The probabilities are calibrated. A raw network tends to be overconfident. A single number, the temperature, was fitted on a separate set of examples and divides the scores before they become probabilities. The aim is that answers given at 80% turn out right about 80% of the time.
What happens on each question
The press release, one question and its answer options. Each option is tagged with a letter: A, B, C…
The 27-billion-parameter network reads all of it in a single pass. It writes nothing.
One linear layer turns the network's final state into a score for each option letter.
The scores are divided by the fitted temperature and turned into probabilities that add up to 1.
Illustration, not a real release.
Every question is its own pass through the network, with no text generated at any point, which is why an analysis takes seconds. There are three kinds of question:
- Yes or no
- Returns the probability of “yes”. Margin pressure, weakening demand and the themes are asked this way.
- Choice
- Returns a probability for each option. Guidance is a choice among seven answers; the likeliest is the one shown.
- Score
- Returns a probability for each level of an ordered scale, and their weighted average. That is why a reading can sit between “mixed” and “somewhat strong”. Results, outlook and caution are scores.
| Base model | Qwen3.8-27B, an open language model |
|---|---|
| Size | 27 billion parameters, 64 layers |
| Output layer | A decision readout with 255 option slots, in place of the layer that generates text |
| Training | Supervised fine-tuning of the whole network on a curated set of 73,000 decision examples |
| Calibration | One temperature (2.21), fitted on a separate set of 3,500 examples |
| Released | 1 October 2026; weights and training code are public (Apache 2.0) |
| Served by | Perplexity's Decisions API, at $0.04 per million input tokens |
Why this model
- It answers instead of writing. Asking a chat model about 2,000 releases gives 2,000 paragraphs that then have to be interpreted. Typed answers can be counted: so many companies raised guidance this quarter, so many describe pressured margins. Every chart on this site is such a count.
- It is consistent. The same text and the same question give the same numbers, give or take a rounding difference, and the model version is pinned. Every release is measured with the same ruler.
- Its probabilities mean something. Because they are calibrated, the site can say that a reading is a close call when it is one, instead of hiding the doubt behind a confident sentence.
- It is good at this kind of reading. On its model card it beats the model it was built from on nine of eleven benchmarks and ties on one, with clear gains on the tasks closest to this one: financial text, and checking a statement against a source.
- It is fast and cheap enough to run for anyone. A release costs about 0.13 cents to read and takes a few seconds, which is what makes an analysis on demand possible.
- It is open. The weights and the training code are public, so the reading can be checked and reproduced without depending on a closed service.
| Benchmark (accuracy) | Qwen3.8-27B | pplx-decider-v1-27b |
|---|---|---|
| FinancialPhraseBankthe sentiment of sentences from financial news | 75.7% | 84.2% |
| RAGTruthwhether a statement is supported by a source text | 61.5% | 88.8% |
| TabFactwhether a statement agrees with a table | 78.6% | 90.6% |
| ContractNLIwhat a contract does and does not say | 80.8% | 80.8% |
| Average of 11 benchmarks | 74.8% | 85.7% |
Figures published by Perplexity on the model card, which has all eleven benchmarks. The API is described in its documentation.
The questions
These are the questions, word for word.
- Guidance
- What does the press release say about the company's financial guidance for future periods?
- Results
- How strong are the financial results reported in the press release?
- Outlook
- How does management describe the business outlook for the coming periods in the press release?
- Caution
- How much uncertainty or caution about the future does management express in the press release?
- Weakening demand
- Does the press release describe weakening customer demand, orders, or bookings?
- Margin pressure
- Does the press release describe declining or pressured profit margins?
The guidance question has seven possible answers:
- Raised
- Guidance for at least one key metric is raised and none is lowered
- Lowered
- Guidance for at least one key metric is lowered and none is raised
- Mixed
- Some guidance is raised and some is lowered
- Reaffirmed
- Previous guidance is reaffirmed or kept unchanged
- New period
- Guidance is given for a new period without saying how it compares with previous guidance
- Withdrawn
- Guidance is withdrawn or suspended
- No guidance
- The release gives no financial guidance
And the theme questions:
- Tariffs
- Does the press release say that tariffs or trade restrictions affect, or are expected to affect, the company's costs, prices, demand or outlook?
- AI
- Does the press release describe artificial intelligence as a driver of the company's demand, revenue or investment?
- Supply chain
- Does the press release say that supply chain disruptions, component shortages or logistics constraints hurt the company's results or outlook?
- Restructuring
- Does the press release report job cuts, a restructuring program or restructuring charges?
Two more questions are asked and not shown. One checks that the filing is a results release at all, which decides whether it is used. The other is a reference question kept from the project's first study.
Two ways it runs
- Live, for one company
- Type a company and the service looks up its Item 2.02 filings on EDGAR at that moment, takes the newest one the model reads as a results release and the one before it, and asks the questions. The steps shown while it runs are the real ones, reported by the server as they happen. Each filing is read once and stored, so asking again for the same release is instant; a new filing is read when the next person asks. It works for companies that report results in a Form 8-K, which leaves out most foreign companies.
- Over time, for the market
- Market trends adds up the same reading over 2,238 releases of 99 companies of the S&P 100 (the list as of December 2020, fixed in advance), from Jan 13, 2021 to Oct 1, 2026. When a company files a second Item 2.02 within 20 days, only the first counts.
The numbers on the charts
A release counts under the guidance answer the model finds most likely. It counts as describing margin pressure, weakening demand, caution or a theme when the model's answer is at or past the midpoint of the scale. Quarters are quarters of publication: a release that comes out in February counts in the first quarter, whatever fiscal period it reports.
Keep in mind
- It reads the press release only: not the financial tables, the conference call or the slides.
- A press release is written by the company. The reading says what the text says, not whether it is true.
- It describes; it does not predict prices or returns, and nothing here is investment advice.
Where it comes from, cost and code
The project began as an experiment: whether the model's reading of a release could predict the stock's return over the following weeks. It could not, and the experiment was closed with that result. What the model did well was read, and that became this site.
Reading the whole history cost about $2.92 in API calls. The pipeline, the questions and the dataset are in the repository.