Before a qualifier starts, Predict can tell your team who's likely to win each match, where you'll probably rank, who might pick you, and your odds of advancing. Here's what's going on under the hood, in plain language, with the real numbers.
Predict is a chain. Each link feeds the next, and every number you see in the app comes out the end of it. First it learns how strong every robot is from the matches it has already played. Then it turns two alliances' strengths into a win probability and a likely score. Then it plays out a whole event a thousand times. Finally it applies the season's advancement rules to see who qualifies in each of those thousand versions.
Every played match nudges each robot's rating for auto, TeleOp, endgame and penalties.
→ a strength per robotAdd up each alliance, account for how noisy FTC scores are, compare the two.
→ win %, score rangePlay the rest of quals, rank teams, pick alliances, run the bracket. Repeat 1,000×.
→ ranks, picks, championsAdvancement points, awards and slots, skipping teams that already qualified.
→ your odds of advancingFTC has a scoring problem for anyone trying to rate teams: you never see one robot's score, only the alliance's. Predict handles that the way the best FRC rating systems do. Each robot carries an expected contribution, in real game points, to whatever alliance it plays on. It's split into four parts, because a robot that's great in auto and weak in endgame is a different partner from one that's the other way around.
After every match, Predict compares what each alliance was expected to score in each part with what it actually scored. The miss is split equally between the two robots, and each robot's rating moves by a fraction of its share. That fraction is the gain.
The gain starts high, so a team's first few matches teach Predict a lot, and it shrinks as evidence builds up. It never drops below a floor, so a team that genuinely improves mid-season still gets noticed.
Each match shows the robot's share of a noisy alliance score (± about 22 pts, as in real FTC). The faint line is the same robot rated with a fixed, cautious gain of 0.33.
To predict a match, Predict adds up each alliance's robots, then adds the penalty points the other alliance tends to give away. That's the expected score.
FTC scores are noisy, though. A ring falls off, a robot tips, a driver misses a park. So each expected score comes with a spread. Fitted on 2024–25, the typical miss is about 14 + 0.2 × expected points: bigger scores have bigger swings. On top of that goes the uncertainty in each robot's rating. A team Predict barely knows gets a wider curve.
Draw both alliances as curves and ask how often red's curve lands to the right of blue's. That share is red's win probability.
Even good models are a little over- or under-confident. Muse added Platt scaling: a tiny correction, fitted on 2024–25, that gently pulls probabilities toward 50% (multiply the log-odds by 0.937, nudge by 0.020). It keeps every prediction in the same order; only a call that was already within half a point of 50% can tip to the other side. It just makes "70%" mean 70%. On the test season it cut calibration error from 0.017 to 0.0105.
σ red 43.2 · σ blue 38.3 · calibrated
Advancing isn't about one match. It depends on your rank, who picks whom, the bracket and the awards. All of those interact, and there's no neat formula for the whole thing. So Predict does what a coach does in their head, only a thousand times over: it imagines the event.
A toy version: 24 robots, 6 qualification matches each with random partners, 2 ranking points per win and the average score as the tiebreaker. Each robot's form for the day is drawn once per run. The real simulator adds bonus RPs, alliance selection, playoffs, awards and the official rules.
Because the simulator can force any pairing, the app answers partner questions. If you're likely to captain: "if we pick team X, our odds become…". If you're likely to be picked: "if captain Y picks us…". When we tested this on real alliances, knowing the actual partner cut the error on "this alliance wins the event" by about 16% compared with general after-quals odds (Brier 0.072 vs 0.086).
A prediction model can look brilliant on the data it learned from. So every setting was tuned on the 2024–25 season, then locked, and the numbers below come from 2025–26, a season it had never seen. Every prediction was made using only matches that had already happened at that moment, exactly as it would be live.
Accuracy isn't the whole story. The Brier score measures how close the probabilities were: 0 is perfect, and a coin flip gets 0.25. It rewards saying 90% when you're sure and 55% when you're not.
Predict improves through small experiments, each judged the same way: tune on 2024–25, test on 2025–26, and compare against the baselines on identical matches. An idea ships only if it helps calibration and doesn't make anything else worse. Plenty of good-sounding ideas don't make it.
A two-number correction to win probabilities. Calibration error on the test season fell from 0.017 to 0.0105, and accuracy barely moved, because favourites almost always stay favourites.
A newcomer on the blue side used to get a slightly different starting rating than one on red, because red's result was counted first. Now both sides see the same starting point.
Giving more of an alliance's surprise to the robot Predict is less sure about (TrueSkill-style). It sounds clever, but it made predictions worse.
Letting high-ranked teams turn down invitations to captain their own alliance. It fit 2024–25 slightly better but predicted real 2025–26 picks clearly worse.
Teams coming back from a long break are hard to predict, because many have rebuilt. Predict now has a mode that grows less certain after a break and learns faster from the next few matches. It does what it says (see below), but after a full refit, the odds of advancing from an event got slightly worse. So it's in the code, switched off, waiting for a better fit.
A 75% chance means one in four times the other thing happens. That's not Predict being wrong; that's what 75% means. Over a season, the calibration chart above is the promise Predict actually makes: when it says 75%, it happens about three times in four.
| Setting | Value | What it does |
|---|---|---|
| First-match gain | 0.65 | How far one match moves a new rating |
| Gain halves after | 10 matches | How quickly a rating settles |
| Gain floor | 0.33 | Keeps late-season improvement visible |
| Playoff weight | 0.5 | Eliminations count half |
| Carry-over | 0.6 / 0.2 | Last season / the one before (as z-scores) |
| Rookie start | z = −0.7 | New teams start a bit below average |
| Season growth | 5% / week | Up to 8 weeks between matches |
| Uncertainty | 150 / 350 pts² | Known team / no history, ×0.45 per match |
| Score spread | 14 + 0.2 × score | Typical miss of an alliance score |
| Event-start extra | 500 pts² | Per robot, for predictions frozen at event start |
| Calibration | a 0.937 · b 0.020 | Platt scaling on win probabilities |
| Captain's pick | τ 18 · rank 3 | How much strength vs. rank drives picks |
| Simulations | 1,000 | Runs per event and starting point |
All fitted on 2024–25 and tested on 2025–26.