Story points explained

Story points are a unit for estimating how big a piece of work is compared with other pieces of work. The number on its own says nothing about hours or days: a 5-point story is one the team considers a bit more than twice the size of a 2 and noticeably smaller than an 8. Everything useful about story points follows from that relative comparison, and it’s the reason they became the most common unit for agile estimation in Scrum and XP teams.

What a story point measures

A story point estimate rolls several things into one number:

  • Effort, or how much work there is. Twenty form fields take longer to build than two.
  • Complexity, or how hard the work is to get right. A tricky algorithm or a fiddly integration can be small in lines of code and still take a lot of thinking.
  • Uncertainty, or how much you don’t know yet. An unfamiliar API, vague requirements or a legacy module nobody has touched in years all make a story riskier, and riskier stories get bigger numbers.

Nobody scores these separately and adds them up. People weigh them in their heads and pick a card, so two developers can arrive at the same 8 for different reasons. One sees a lot of repetitive work, the other sees a risky database migration. When the reasons differ that much, a short conversation is worthwhile, and that conversation is exactly what planning poker is designed to trigger.

Why not just estimate in hours?

Hours feel more concrete, and people outside the team usually understand them better. Relative sizes avoid a few problems that hour estimates run into.

People are much better at comparing than at measuring. Ask someone how tall a building is and they’ll hesitate; ask whether it’s taller than the one next door and they’ll answer immediately. Software is similar. “Is this bigger than the login page?” is an easier question than “how many hours will this take?”, and the answers tend to be more consistent.

Hours also depend on who does the work. A senior engineer who wrote the payment module might finish a change in three hours, while a new hire might need two days for the same change. Story points describe the work rather than the person, which lets the whole team agree on one number.

Then there’s the way hour estimates harden into commitments. Once “12 hours” is written on a ticket it tends to become a deadline, and people feel judged when they miss it. A points estimate reads more obviously as a size, so it’s easier to talk openly about risk.

Finally, hours leave out everything that isn’t coding: code review, meetings, waiting on another team, fixing a flaky CI pipeline. Velocity measured in points absorbs all of that, because it’s based on what the team actually finished.

Pick a reference story

Relative estimation needs something to compare against. Before your first session, pick one or two stories the team has already finished and remembers well, and give them values.

For example, take a small, well-understood story such as “add a phone number field to the user profile and save it”, and call it a 2. Then find a medium one, perhaps a new report page with a couple of filters, and agree it’s roughly a 5. From then on, every new story gets compared with those two: bigger than the profile field? About the same as the report page?

Write the reference stories down where the team can see them, at the top of the backlog or pinned in the team channel. After a few months the scale lives in people’s heads, but new joiners will still need the examples.

Velocity: what the points are for

Velocity is the total number of story points a team completes in a sprint. If you finish stories worth 5, 3, 8, 2 and 5, that sprint’s velocity is 23. Stories that are 90% done at the end of the sprint contribute nothing, which feels harsh but keeps the numbers honest.

A single sprint tells you very little. After three or four you’ll start to see a range, something like 21 to 27, and that range is what you plan with. It tells the team roughly how much to take into the next sprint, and it lets the product owner make a forecast: with 120 points left on the release backlog and a velocity of around 24, you’re looking at about five sprints.

For a team whose membership and kind of work stay fairly stable, velocity tends to settle down after a few sprints. Individual estimates are off in both directions, some 5s turn out to be 3s and others 8s, and over a whole sprint many of those errors cancel out. It won’t settle if the team changes a lot, if a sprint is dominated by unplanned work, or if the scale quietly drifts.

Where velocity fits into refinement and sprint planning is covered on the Scrum poker page.

Myths that cause trouble

“One point equals one day”

Sooner or later someone will want to convert points to hours. Once a point means a fixed number of hours, you’ve got hour estimates again, with an extra conversion step and all the problems above. If someone needs a date, use velocity to forecast in sprints.

“We can compare teams by velocity”

Team A has a velocity of 40 and Team B has 20. It would be natural to conclude that Team A is twice as productive, and it would be wrong: each team has its own reference stories and its own scale, and Team A’s 5 might be Team B’s 2. When management compares velocities across teams anyway, teams respond by inflating estimates. Velocity rises, output doesn’t change, and the numbers lose their value for planning.

“Higher velocity is always better”

A steady velocity is more useful than a rising one, because it makes forecasts reliable. When velocity jumps suddenly, check whether the scale has drifted before celebrating.

“Every story needs an exact number”

Story points are deliberately rough. Most teams use a Fibonacci-style scale where the gaps widen as the numbers grow, so there’s no 14 to argue about when the choice is between 13 and 20.

How to get started

If your team has never used story points, a low-pressure way to begin:

  1. Agree on one small and one medium reference story, as described above.
  2. Use the standard deck: 0, ½, 1, 2, 3, 5, 8, 13, 20, 40, 100, with “?” and a coffee card.
  3. In one session, estimate the top of the backlog with planning poker so nobody anchors the others; a free online room with no sign-up is enough. Two sprints’ worth of stories is plenty; there’s no need to size the whole backlog.
  4. For the first few sprints, don’t worry about accuracy. The scale needs time to settle.
  5. Track velocity for three or four sprints before you lean on it for planning.
  6. After a couple of months, look at your reference stories again. If people keep saying “that’s bigger than our old 5”, recalibrate.

The first estimates will be off, sometimes by a lot, and the misses are useful. When a 3 turns into a week of work, talk about why in the retrospective; the next similar story will get a better number.