What is Score Estimation and Confidence Level
Learn how MockGarden turns your Mock Test practice into per-Section Band Score estimates, and how Section Confidence shows how reliable each estimate is.
When you practice on MockGarden, every completed Mock Test feeds into Analytics — the place where you see how you are performing across Reading, Listening, Writing, and Speaking. Two numbers work together there: your Score Estimation for each Section, and the Confidence Level that tells you how much to trust it.
This guide explains how both are calculated, what practice counts, and what you can do when a Confidence Level is still low.
What Score Estimation is
A Score Estimation is MockGarden's best current guess at your Band Score for one TOEFL Section, based on the practice you have already completed. It is shown separately for Reading, Listening, Writing, and Speaking.
It is not a single test result. Instead, MockGarden looks across your attempts in the time range you select — last week, last month, last three months, or all time — and combines performance from every Task Type in that Section.
Each Section on the TOEFL is made up of several Task Types. For example, Reading includes Fill in the Letters, Daily Life, and Academic. Listening includes Listen and Choose a Response, Conversation, Announcement, and Academic Talk. Writing and Speaking follow the same pattern.
Every Task Type carries a weight that reflects how much of the full Section it represents on test day. MockGarden uses those weights so that stronger performance on a high-impact task moves your estimate more than the same performance on a smaller one.
For how official section bands and the overall score are calculated, see How Is the TOEFL Scored?. For task formats by section, use the TOEFL 2026 Learning Guide.
How Score Estimation is calculated
The calculation happens in three steps for each Section.
Step 1: Average your performance by Task Type
For each Task Type in the Section, MockGarden finds every completed attempt in your selected time window and calculates your average percentage score for that task.
A Task Practice focused on Daily Life reading counts. So does the Daily Life portion of a Reading Section Test or a Full Mock Test. If you have not attempted a Task Type yet, it is marked as missing and does not contribute to the average.
Step 2: Combine Task Types with their weights
MockGarden then builds a weighted average across the Task Types that have data:
Section percentage = sum of (task average × task weight) ÷ sum of weights for tasks with data
Only Task Types you have actually practiced are included. If you have done Cloze and Academic reading but not Daily Life, MockGarden estimates Reading from the tasks you have, rebalanced by their weights. Missing tasks are not treated as zero — they simply wait until you practice them.
Here is a simplified Reading example:
| Focus | Before | After |
|---|---|---|
| Fill in the Letters (50%) | Average 72% across 8 attempts | Contributes 36 points to the weighted total |
| Daily Life (20%) | Average 80% across 3 attempts | Contributes 16 points to the weighted total |
| Academic (30%) | Average 68% across 5 attempts | Contributes 20.4 points to the weighted total |
| Estimated Reading percentage | — | 72.4% → converted to Band Score |
Step 3: Convert to Raw Score and Band Score
The weighted percentage is scaled to the Section's raw score range using our own internal algorithm and then mapped to MockGarden's 0–6 half-band scale using the same conversion tables used when Judging a Mock Test.
That final number is your Score Estimation for the Section.
What Confidence Level is
A Score Estimation answers where you might be. Confidence Level answers how much you should trust that number right now.
Confidence Level is scored 0–100 for each Section independently. A higher score means the estimate is based on enough recent, stable, exam-like practice. A lower score means the estimate may still move a lot as you complete more Mock Tests.
MockGarden also maps the score to a label:
- Very High (90–100)
- High (75–89)
- Medium (50–74)
- Low (25–49)
- Very Low (0–24)
When Confidence is Very High, the explanation reads: "This estimate is very reliable because it is based on strong, recent, stable, and realistic test data." At Very Low, MockGarden tells you the estimate is still early and may change significantly after more practice.
How Confidence Level is calculated
Confidence Level combines four factors. Each factor is scored 0–100, then blended with fixed weights:
| Factor | Weight | What it measures |
|---|---|---|
| Practice coverage | 40% | How many attempts you have across Task Types |
| Score consistency | 30% | How stable your recent scores are |
| Recent activity | 20% | How fresh your latest practice is |
| Exam-like practice | 10% | How much data comes from Section Tests and Full Mock Tests |
The final Confidence Level is:
Confidence = round(0.40 × coverage + 0.30 × consistency + 0.20 × recency + 0.10 × realism)
Practice coverage
Practice coverage rewards breadth. More practice across more Task Types raises coverage. A Section where you have only done one isolated Task Practice a few times will score lower here than a Section where you have repeated every task type several times.
Score consistency
Score consistency looks at your most recent attempts for each Task Type and measures how much your percentages vary.
If your last several Cloze attempts cluster around 75–78%, consistency is high. If they swing from 55% to 90%, consistency is lower. With only one attempt on a task, consistency starts conservatively until more data arrives.
MockGarden uses the standard deviation of those recent scores and applies a small multiplier when you have fewer than seven attempts, so early estimates do not look artificially stable.
Recent activity
Recent activity checks how fresh your practice is. Regular practice keeps recency high; long gaps pull it down even if your historical average is strong.
Exam-like practice
Not all practice carries the same exam weight. Task Practice is valuable for building skill on one Task Type, but a Section Test or Full Mock Test better reflects real timing, section flow, and mixed tasks.
More Section Tests and Full Mock Tests raise realism faster than Task Practice alone.
What MockGarden suggests when Confidence is low
When Confidence is not yet high, Analytics also surfaces a main blocker and a next action — for example, missing data for a specific Task Type, scores that vary a lot, outdated practice, or mostly isolated Task Practice.
Common patterns:
- Missing Task Types → complete a Section Test that covers the full Section
- Low practice coverage → complete more attempts across Task Types
- Unstable recent scores → complete a few more attempts to confirm your level
- Stale data → take a fresh Section Test or Task Practice
- Low exam-like practice → prioritize a Section Test or Full Mock Test
These are guidance, not gates. You can always see your current Score Estimation; Confidence simply tells you how much weight to give it today.
A practical way to read both numbers together
Use Score Estimation and Confidence Level as a pair:
- Higher Confidence → the Band Score estimate is a solid planning number for study decisions and mock scheduling.
- Lower Confidence → treat the estimate as directionally useful, but expect it to move as you fill gaps, stabilize scores, or add recent Section Tests.
A Student who has done ten Task Practice sessions on Cloze alone might show a Reading estimate with Low Confidence because Daily Life and Academic are missing and realism is still task-heavy. After one Reading Section Test covering all three Task Types, both the estimate and Confidence usually become more meaningful.
Does a bad score on one Mock Test permanently lower my estimate?
No. Score Estimation uses averages across your attempts in the selected time range. One weak session affects the average, but later strong sessions balance it out — especially when your scores become more consistent. You can also change the time range to focus on more recent attempts and see a clearer picture of your current level.
Why can my Confidence Level differ between Reading and Speaking?
Confidence is calculated per Section from that Section's Task Type history. You might have many recent Reading Section Tests but only a few Speaking Task Practice sessions, so Reading Confidence can be High while Speaking Confidence is still Medium or Low.
Does changing the Analytics time range change Confidence?
No. Last week, last month, last three months, only affects your score estimate. Confidence level, however, is generally increasing over consistent practice. As we have more data, our estimate of your score will be more accurate over time.