We Tracked 30 Days of Daily Football Picks – Here’s What the Numbers Showed

Every morning, football prediction feeds offer the same temptation: a short list of matches, a confidence label, and the promise that today’s selections are different from yesterday’s. The only sensible way to judge that promise is to keep the scoreboard open after the final whistle.

We reviewed the latest complete 30-day window in NerdyTips’ public performance log, covering 10 August to 8 September 2026. The period contained 148 high-confidence selections, described on the platform as “bankers”, and 9,019 entries in the wider all-matches record. The headline result is respectable but less dramatic than a winning screenshot suggests: 102 of the 148 banker picks were correct, a weighted hit rate of 68.9%. Across the broader pool, 6,025 of 9,019 predictions were correct, or 66.8%.

We Tracked 30 Days of Daily Football Picks - Here's What the Numbers Showed

That two-point gap is the most important number in the review. It suggests that the highest-rated selections performed better than the full list, but only modestly. It does not, by itself, prove that the picks produced a profit. Once odds, bookmaker margin, stake size and the timing of the price are added, the picture becomes more demanding.

What exactly was tracked?

The exercise used the provider’s own public data rather than a hand-picked list of results. Its open data repository contains a daily progress file with banker counts, daily percentages and the corresponding all-matches figures. A detailed workbook adds fixtures, markets, trust scores, recorded odds and final scores.

To avoid a misleading average, we added the daily counts and weighted each day’s percentage by the number of selections. A day with one pick therefore counted less than a Saturday with 17. The simple average of the daily banker percentages on the 27 active days was 61.1%, while the pick-level figure was 68.9%. The first measure weights days equally; the second answers the practical question of how many selections were right overall.

The public log recorded no banker pick on three dates. On the other 27, volume ranged from one selection to 17, averaging about 5.5 picks on an active day. A perfect one-from-one day carries far less evidence than a 13-from-17 performance on a crowded fixture list.

The word “banker” also needs care. The service reserves it for higher-rated calls. We treated the label as a selection rule, not a literal probability: a 10/10 score is not the same as a published 100% chance, and the files do not define a conversion from trust score to probability.

The headline: a solid hit rate, not a runaway edge

The 68.9% banker result means 46 of the 148 selections missed. That is a useful record for a short-term review, especially because football produces plenty of low-scoring matches, late goals and results that turn on a single incident. It is also a result that becomes less spectacular once the comparison group is included.

But the all-matches record was 66.8%, and the high-confidence subset was 2.1 points higher. The filter uncovered a slight enhancement in the interior of a very large class, but not a very different level of certainty. That’s an observable, but still relatively small, opening that can be filled as more results are added.

The same applies to the first half of the window as it does to the second half. The banker list for 10-24 August was 68.3% winning, with 43 wins out of 63 picks. It won 59 times on 85 entries (69.4%) from 25 August to 8 September. This was slightly improved during the latter period, but only by a little over one per cent. It is not as if there was a leap to a new level of performance.

The daily results were far more unpredictable than the 30-day results. 5 days were active, with 5 of those days ending at 100%. On the other side, three active days saw finishes of 0%, while 23 August was 3 of 8 for 37.5%. There is no reason why a record cannot have both extremes, and they are just added together to get the aggregate.

A line of green ticks is convincing but can be easily broken. The rate could fluctuate by a few points with 148 picks. The record covers the month and can’t determine what occurs in a season.

Weekends supplied most of the evidence.

It also had to do with a change in the sample that occurred during the calendar. The 88 selections and 63 winning dates for the weekend give an average hit rate of 71.6%. The win percentage on weekdays was 65.0% with 60 picks and 39 wins. Selections on the weekend provided 59.5% of the selections and the highest percentage in this period.

That doesn’t mean that the model is stronger on Saturdays and Sundays per se. There are just 8 dates available on the weekend in the 30-day period, but 22 dates during the week, and the number of fixtures varies widely by day of the week. A greater number of games will generate more confidence in decisions, but also more instances where a model can be incorrect. A longer audit is a good indicator of the weekend split, but not a bet.

The market was broader than a string of home wins.

The detailed workbook reveals that it wasn’t a single bet to all the banks. The largest number of bets was placed on home wins, at 53, and the next most popular choice was away wins, at 15. There were also 12 away double-chance calls, and a home double-chance call, as well as a variety of goal markets.

There were 18 times when less than 2.5 goals were scored, 13 times when more than 2.5 goals were scored, 10 times when more than 3.5 goals were scored, and 4 times when 3.5 or fewer goals were scored. Six more were added by the “both teams to score” picks. Those were the only calls that were not flat, with the rest being goal handicaps and team-scoring markets. That is, about one-third were totals and two-thirds were in result, protection, and scoring.

The combination of these accuracy measures is important because one accuracy number can mask varying levels of difficulty. The risks and prices of a home win, an under 2.5 line and an away double chance are not the same. A fair assessment should report on hit rate and return by market, league, and odds band, not just return on a single blended return.

The public listing at https://nerdytips.com/bet-of-the-day makes that distinction visible because it presents the match, suggested market, and price rather than only a win-or-loss badge. A prediction that wins at 1.18 is not economically equivalent to one that wins at 1.70.

Low odds change the meaning of a win.

Prices recorded underline the point. Of the 148 banker picks in the detailed 30-day sample, 114 of them were under 1.50. That total included 46 prices between 1.20 and 1.29 and 63 between 1.30 and 1.49. 34 were in the range of 1.50 to 1.99, and five were less than 1.20. The minimum price given was approximately 1.14, the maximum was about 1.74, and the simple mean was approximately 1.38.

A flat-stake bettor must win about 72.3% of their bets for the decimal odds to pay out before any account limits, price changes, or other frictional factors. The observed rate is 68.9%, which is below that. The average price is an approximation of the average loss per 100 units staked, and if the individual bets had the same price, each bet would lose about 4.6 units. This is an inference based on the average price and not an exact ledger of bets.

The difference is key. Accuracy is the question of whether or not a prediction was correct; return is the question of whether or not it paid off for the player. A gambler might win a lot of his wagers and still lose money when there are short odds. If someone wins 1.20, he makes only 0.20, and if he loses 1, he loses the profit from 5 wins. A 1.70 winner will add 0.70 units and will tip over faster.

Nor are the odds themselves probabilities. They include the bookmaker’s profit and can change throughout the tip’s submission and until kick-off. Hence, a serious 30-day report ought to indicate if it was opened odds, first price, best price across books, or closing line. If that is not included, it is hard to replicate “profit,” and accuracy is more easily compared with “yield.

In this review, we have not included any claimed profit figure. Progress file reports out percentages; the detailed workbook provides prices and market areas. The progress file was used for the odds and market description, and the workbook was used for the aggregate accuracy, since they did not have the same number of wider pools. The only reason to be precise is that you shouldn’t throw away the data.

A confidence label is not a probability forecast.

The most widely made error when reading a tipster record is to try and interpret the confidence level as a statistical prediction. The 9/10 or 10/10 label may indicate preferred choices, but the log does not necessarily have to indicate that each 10/10 pick had a 90% or 100% probability. The 0% days are a very plain reminder that a high ranking doesn’t eliminate uncertainty.

The data would have to be clustered into confidence bands to test the score. Analysts were able to compare calls that received a rating from 9.0 to 9.4, 9.5 to 9.9, and 10.0, and then do the same over a much longer time. A group is calibrated if it has an 80% win rate around 80%. A model that wins 6 out of 10 games could have higher confidence in its picks than its record might suggest, but it may still have flaws.

Forecasting research has established that distinction. Conor Walsh and Alok Joshi wrote an article this past year (2023) about calibration versus accuracy in machine learning sports betting. Their experiment focused on NBA data instead of football, and so their numerical findings can’t be taken over here. The overall lesson does stand: A model should be evaluated based on its probability of success, not solely its success of having the favorite outcome occur.

A football prediction can thus come in handy in two ways. It may differentiate – it can recognize which outcomes are more likely to be achieved than other candidates. It can also be calibrated, so that it is accurate in reflecting the percentage of success. The 30-day log will provide evidence for the first question. It is not a complete response to the second question.

What the losing days reveal

The losses are not something to be embarrassed by or removed from the sample. They give an account of its boundaries. The poorest cluster occurred on 23 August when a fifth of eight banker calls went astray. It was another challenging day on 29 August (3 wins out of 6), and 4 September was 4 wins out of 8. If the daily list is more than one or two matches long, then the record can move rapidly.

They also demonstrate that the phrase “daily” is ambiguous. If you imagine a reader, he might picture one best pick coming to him each morning. The public banker log could have multiple selections, and the quantity was based on the fixture card and the model’s threshold. There can be one high-trust call or none on a quiet day. There can be up to 13 or 17 on a full weekend. The record is now a series of selections, rather than 30 separate predictions.

This is important for accumulators. For the five individual games, the probability of each one winning is 70%, and assuming that these are independent, the probability of all five winners is 0.70^5, or approximately 16.8%. Football matches are not independent, and the actual number may be more or less, but the multiplication is a good illustration of the basic problem. So, if you can achieve a good individual hit rate, but get too many legs together, it can be a poor accumulator record.

The month also offers little insight into correlation. Losses can come in en masse when picks share a league and/or team news, or market assumptions. A decent audit ought to identify related choices and state the outcome of taking a single wager and any “Multi-leg” slip that was advertised.

How a fairer 90-day test would work

The next review should be more specific and lengthier. All picks must have a pre-kick-off time, market, price and settlement rule. The sample should contain all the qualifying selections and indicate the absence of a banker. The results should be broken down by market and odds band, not reduced to one per cent point.

The benchmark should be the market itself. Seems like a solid 72%, but isn’t an edge when the price is taken into account. A fair probability estimate based on the closing market, and the price at which a reader could make a bet, provides a much more difficult and worthwhile test.

The report should then be published with three different outcomes: hit rate, flat stake return, and calibration. Hit rate — indicates to the reader how many of the calls landed. The return column indicates if the odds were applied to the misses. Calibration indicates if the confidence labels are appropriate probabilities. These measures address different questions and should not be combined into a single “accuracy” badge.

Publish uncertainty in the report too! This is an observation, not an ability level: 68.9% of the 148 picks were correct. The more results that are there, the more stable the estimate will be, and the fewer, the more sensitive to one late goal. It’s a nice change from a streak graphic, but it is more beneficial to readers.

The verdict after 30 days

Numbers provide evidence for an informed conclusion. The high-confidence group outperformed the provider’s larger all-match group (68.9% vs. 66.8%), and the latter half of the window was slightly superior to the first half. Weekend picks were superior to weekday picks, but there was a skewed sample size. The model didn’t select a banker on three days, and accuracy fluctuated between 0% and 100% for each day depending on the fixtures played.

While these are great pluses, they’re not necessarily signs of a profitable system. Most prices have been short, and the average price has been around 1.38, which suggests a higher-than-observed hit rate in the simple break-even rate. The public record shows activity and provides information for readers to continue the audit. It does not mean to justify the use of a confidence label as an assurance or a green month as a promise of future returns.

The nifty outcome of the 30-day tracking project is not a magic percentage, then. It’s a more straightforward set of questions. After price, which markets were profitable? Can you keep your confidence on par for hundreds of picks? Is there anything different about weekends, or just more work? How many times does one of the assumptions support multiple losses? Until the answers to those questions are known, the straight answer is that the picks were slightly more accurate than the general list, and the price profile was even more challenging to be accurate on in order to break even.

That’s not as flashy an ending as “the algorithm found the winners”. It is also the one that the numbers are able to sustain.

Popular on OTW Right Now!

Add a Comment

Your email address will not be published. Required fields are marked *