Control Point·Predict, explained

How Control Point Predict works

Before a qualifier starts, Predict can tell your team who's likely to win each match, where you'll probably rank, who might pick you, and your odds of advancing. Here's what's going on under the hood, in plain language, with the real numbers.

Q-27 · example qualifierLive ratings
Red alliance 66% expected 145 pts
80% range 90–200
Blue alliance 34% expected 120 pts
80% range 71–169
A 66% favourite loses about one match in three. That's the point: the number is a probability, not a promise.
01The big picture

Four steps, from match history to advancement odds

Predict is a chain. Each link feeds the next, and every number you see in the app comes out the end of it. First it learns how strong every robot is from the matches it has already played. Then it turns two alliances' strengths into a win probability and a likely score. Then it plays out a whole event a thousand times. Finally it applies the season's advancement rules to see who qualifies in each of those thousand versions.

Figure 1
The Predict pipeline · tap a step
The data comes from FTC Scout (every match since 2022) and the official FIRST Events API (schedules, ranks, alliances and advancement lists). During a live event the server keeps both in sync in the background.
02Step one

A rating is a guess at how many points a robot adds

FTC has a scoring problem for anyone trying to rate teams: you never see one robot's score, only the alliance's. Predict handles that the way the best FRC rating systems do. Each robot carries an expected contribution, in real game points, to whatever alliance it plays on. It's split into four parts, because a robot that's great in auto and weak in endgame is a different partner from one that's the other way around.

Figure 2
One robot's rating (example)
Auto
18.4
points in the 30 s autonomous period
TeleOp
41.2
driver-controlled, not counting endgame
Endgame
12.9
parking, climbing, the last 20 s
Penalties
3.1
points it tends to give the other alliance
Strength (the non-penalty part) is 18.4 + 41.2 + 12.9 = 72.5 points. Penalties are kept apart because they score for the opponent.

How a rating moves after a match

After every match, Predict compares what each alliance was expected to score in each part with what it actually scored. The miss is split equally between the two robots, and each robot's rating moves by a fraction of its share. That fraction is the gain.

Red expected 145, scored 165 → miss = +20
each robot's share = 20 ÷ 2 = +10
robot on its 3rd match: gain = 0.65 ÷ (1 + 3/10) = 0.50
rating moves by 0.50 × 10 = +5.0 points

The gain starts high, so a team's first few matches teach Predict a lot, and it shrinks as evidence builds up. It never drops below a floor, so a team that genuinely improves mid-season still gets noticed.

Figure 3
Gain vs. matches played this season · hover the curve
gain = max(0.33, 0.65 ÷ (1 + n/10)). Playoff matches count half, because teams play differently in eliminations (defence, strategy, different partners).
Interactive
Watch a rating learn
Where Predict starts
0matches played
55Predict's rating
0.65current gain

Each match shows the robot's share of a noisy alliance score (± about 22 pts, as in real FTC). The faint line is the same robot rated with a fixed, cautious gain of 0.33.

Fast at first, steady later. A high early gain gets a new team close to its real level within a handful of matches. The floor keeps the rating moving if the robot really does change.

Three more things a rating knows

  • Robots get better during a season. A rating grows by 5% for every week since the team last played, for up to 8 weeks. Two months between a team's first qualifier and its next one is a lot of build time.
  • How sure it is. Each rating carries an uncertainty. A team Predict knows from last season starts at 150 points², a team with no history at 350. Each match shrinks that to 45% of what it was, so after a handful of matches Predict is fairly confident.
  • Last season, a little. Games change every year, so points don't carry over. Instead, a team's final rank within its season becomes a z-score (how many standard deviations above or below average it was). That carries into the new season at 60% weight, plus 20% of the season before. Teams with no history start a bit below average (z = −0.7), which matches how rookies actually do.
Figure 4 · interactive
Where a team starts a new season
Starting z = 0.6 × last season + 0.2 × the season before. The z-score then becomes points using the spread of scores seen so far in the new season, so the start adapts to whatever this year's game rewards.
03Step two

From two alliances to one probability

To predict a match, Predict adds up each alliance's robots, then adds the penalty points the other alliance tends to give away. That's the expected score.

FTC scores are noisy, though. A ring falls off, a robot tips, a driver misses a park. So each expected score comes with a spread. Fitted on 2024–25, the typical miss is about 14 + 0.2 × expected points: bigger scores have bigger swings. On top of that goes the uncertainty in each robot's rating. A team Predict barely knows gets a wider curve.

Draw both alliances as curves and ask how often red's curve lands to the right of blue's. That share is red's win probability.

Figure 5
Two score curves, one win probability
Red · expected 145Blue · expected 120Shaded: 80% score range
The 25-point gap is real, but the curves overlap a lot, so red wins about two times in three. The shaded bands are the 80% ranges shown in the app; in testing, real scores landed inside them 82% of the time.

A final calibration touch

Even good models are a little over- or under-confident. Muse added Platt scaling: a tiny correction, fitted on 2024–25, that gently pulls probabilities toward 50% (multiply the log-odds by 0.937, nudge by 0.020). It keeps every prediction in the same order; only a call that was already within half a point of 50% can tip to the other side. It just makes "70%" mean 70%. On the test season it cut calibration error from 0.017 to 0.0105.

Interactive
Win-probability playground
How well Predict knows these robots
66%chance red wins

σ red 43.2 · σ blue 38.3 · calibrated

This uses the same formula as the app: σ = 14 + 0.2 × expected, plus rating uncertainty, then Platt scaling. "Well known" adds 10 points² per robot, "First match" 350, and "Before the event" 150 plus the extra 500 that event-start predictions carry, because a team's level moves during an event.
04Step three

Playing the whole event a thousand times

Advancing isn't about one match. It depends on your rank, who picks whom, the bracket and the awards. All of those interact, and there's no neat formula for the whole thing. So Predict does what a coach does in their head, only a thousand times over: it imagines the event.

Figure 6
One simulated run
A thousand runs give a thousand outcomes. If your team advances in 412 of them, your odds are 41%. The same runs give your likely rank range, your chance of being a captain or a pick, and your chance of winning.
Interactive
Run a 24-team qualifier 1,000 times
Your robot
0events run
–seeded top 4
–typical rank

A toy version: 24 robots, 6 qualification matches each with random partners, 2 ranking points per win and the average score as the tiebreaker. Each robot's form for the day is drawn once per run. The real simulator adds bonus RPs, alliance selection, playoffs, awards and the official rules.

The spread is the point. Even a top-3 robot sometimes lands outside the top 4 because of the schedule, its partners and a bad day; an underdog occasionally sneaks in. Predict reports exactly these chances instead of a single guess.

The details that make it realistic

  • A robot's day is drawn once per run. If a team is having a great day in one simulated event, it's great all day long, not great in one match and awful in the next. That's how real events feel.
  • Played matches stay played. Halfway through quals, Predict keeps the real results and only simulates what's left. The same goes for real alliances and real playoff results once they happen.
  • Ranking points follow the real rules. 2025–26 gives 3 for a win and 1 for a tie, plus movement, goal and pattern bonus RPs, whose odds rise with an alliance's score. The tiebreaker is average non-penalty score. Checked against official data, the rules match 99.7% of real rankings.
  • Captains pick like real captains. Fitted on 2,072 real picks, captains favour strong robots and, a bit, high-ranked ones. The real first pick was in the model's top three about 70% of the time.
  • Real brackets. Double elimination for 4, 6 or 8 alliances with the grand-final rematch, or a best-of-3 when there are only 2.
  • Awards matter. In 2025–26, Inspire alone is worth 60 advancement points. Predict estimates award chances from a team's award history and robot strength, and they're a big part of why it beats a "top-ranked teams advance" guess.
Figure 7
2025–26 advancement points, at a glance
All of these were read off official FIRST data and checked against every published total (1,786 of 1,786 match). Only a team's highest award counts. The quals points come from a bell-curve formula on rank, so a mid-table team still earns a few.

"What if we pick them?"

Because the simulator can force any pairing, the app answers partner questions. If you're likely to captain: "if we pick team X, our odds become…". If you're likely to be picked: "if captain Y picks us…". When we tested this on real alliances, knowing the actual partner cut the error on "this alliance wins the event" by about 16% compared with general after-quals odds (Brier 0.072 vs 0.086).

05The scoreboard

How good is it, really?

A prediction model can look brilliant on the data it learned from. So every setting was tuned on the 2024–25 season, then locked, and the numbers below come from 2025–26, a season it had never seen. Every prediction was made using only matches that had already happened at that moment, exactly as it would be live.

72.7%of matches called right with live ratings
82%of real scores inside the 80% range (honest ranges)
0.064advancement Brier after alliance selection (lower is better)
37,395test matches · 508 advancing events

Accuracy isn't the whole story. The Brier score measures how close the probabilities were: 0 is perfect, and a coin flip gets 0.25. It rewards saying 90% when you're sure and 55% when you're not.

Figure 8
Match predictions vs. simple baselines (2025–26, Brier, lower is better)
OPR from a team's last event is what many teams use today. Live ratings beat it by a wide margin. The season-average baseline is surprisingly strong, but it's only available once a team has played this season, and it has no notion of uncertainty, partners or awards.
Figure 9
It learns fast: accuracy by matches already played
Predict, liveLast-event OPR
Fewest prior matches among the four robots in the match. Early in a season, carrying last season over (as a z-score) is what keeps Predict well ahead of OPR, which has almost nothing to go on yet.
Figure 10
Does "70%" mean 70%? Calibration on the test season
Match win odds (37,395 matches)Advancement odds before the eventPerfect calibration
Each dot groups predictions of about the same probability and shows how often that outcome really happened. Dots on the diagonal mean the numbers can be taken at face value. Hover a dot for its numbers.
Figure 11
Advancement odds sharpen as the event unfolds (Brier, lower is better)
Even before a single match is played, Predict beats the naive guess that can only be made after quals. Before the event, its rank estimates are off by about 5 places on average, and the 10–90% rank range catches the real rank 86% of the time.
06Lab notes

What we tried, and what made the cut

Predict improves through small experiments, each judged the same way: tune on 2024–25, test on 2025–26, and compare against the baselines on identical matches. An idea ships only if it helps calibration and doesn't make anything else worse. Plenty of good-sounding ideas don't make it.

ShippedPlatt calibration

A two-number correction to win probabilities. Calibration error on the test season fell from 0.017 to 0.0105, and accuracy barely moved, because favourites almost always stay favourites.

Fitted by maximum likelihood on 2024–25 live predictions: a = 0.937, b = 0.020. On 2025–26, live Brier went from 0.1791 to 0.1789 and pre-event accuracy from 68.65% to 68.78%. The correction keeps predictions in the same order. Its small offset can tip a raw call between about 49.5% and 50% to the other side, which is where that small accuracy change comes from.
ShippedFair starting ratings

A newcomer on the blue side used to get a slightly different starting rating than one on red, because red's result was counted first. Now both sides see the same starting point.

It affects about one match per season replay, so every test metric moved by 0.0002 or less. It's a correctness fix: swapping red and blue now gives mirrored ratings, and a blue newcomer's start no longer depends on red's score.
RejectedSmarter credit splitting

Giving more of an alliance's surprise to the robot Predict is less sure about (TrueSkill-style). It sounds clever, but it made predictions worse.

On 2025–26, live accuracy dropped 0.25 points and Brier rose from 0.1791 to 0.1806, with worse calibration. The other settings were tuned for equal sharing, and changing the split without a full retune hurt.
RejectedAlliance declines

Letting high-ranked teams turn down invitations to captain their own alliance. It fit 2024–25 slightly better but predicted real 2025–26 picks clearly worse.

Pick log-likelihood on 2,540 real 2025–26 picks: −5,080 with declines vs −4,952 without. Advancement Brier slipped too (0.1287 → 0.1290 before the event, 0.0855 → 0.0866 after quals). And in playoffs, favourites actually win more often than predicted, so stacked alliances aren't being over-rated.
Built, switched offRobot rebuild detection

Teams coming back from a long break are hard to predict, because many have rebuilt. Predict now has a mode that grows less certain after a break and learns faster from the next few matches. It does what it says (see below), but after a full refit, the odds of advancing from an event got slightly worse. So it's in the code, switched off, waiting for a better fit.

After a break of 4+ weeks, uncertainty grows by 200 points² per week away and the learning rate resets as if the robot had played 4 matches. On 2025–26, the error on the 2nd–3rd match after 8+ weeks off fell from 33.6 to 32.1 points, and overall score error from 27.95 to 27.88. Win odds were unchanged, but advancement Brier went from 0.1286 to 0.1291, so it stays off.
Figure 12
Why rebuilds are hard: score error after a break (2024–25)
Alliance score error (MAE, points) on the tuning season. On a robot's first match after two months off, the error is more than 2½ times the usual. A rebuilt robot only shows itself once it plays, so the error stays high for a few matches.
07Fine print

What Predict can't know

  • A new robot until it plays. If you rebuilt over the break, Predict still thinks you're the old robot until your first match.
  • The 2025–26 bonus RPs were fitted on the season's earliest matches, because the bonuses didn't exist before. A refit with the full season is waiting on official data access.
  • Three-team championship alliances aren't modelled yet.
  • Which awards exist at an event is taken as known; who wins them is predicted.
  • Next season's rules. Until FIRST publishes the 2026–27 rules, the 2025–26 points system is assumed.
Treat it like a scouting report from someone who has watched every match ever played. It's very good, but it's still a forecast.

A 75% chance means one in four times the other thing happens. That's not Predict being wrong; that's what 75% means. Over a season, the calibration chart above is the promise Predict actually makes: when it says 75%, it happens about three times in four.

For the curious

The settings, in one table

SettingValueWhat it does
First-match gain0.65How far one match moves a new rating
Gain halves after10 matchesHow quickly a rating settles
Gain floor0.33Keeps late-season improvement visible
Playoff weight0.5Eliminations count half
Carry-over0.6 / 0.2Last season / the one before (as z-scores)
Rookie startz = −0.7New teams start a bit below average
Season growth5% / weekUp to 8 weeks between matches
Uncertainty150 / 350 pts²Known team / no history, ×0.45 per match
Score spread14 + 0.2 × scoreTypical miss of an alliance score
Event-start extra500 pts²Per robot, for predictions frozen at event start
Calibrationa 0.937 · b 0.020Platt scaling on win probabilities
Captain's pickτ 18 · rank 3How much strength vs. rank drives picks
Simulations1,000Runs per event and starting point

All fitted on 2024–25 and tested on 2025–26.