Data Interpretation

TSI Math Study Guide

Data Interpretation

Data can be presented in tables, graphs, and other visual displays. On the TSI, you may need to identify the type of data being collected, choose an appropriate graph, read information from a display, compare datasets, and recognize patterns such as trends, clusters, and outliers.

The most important skill isn’t reading individual numbers. It’s understanding what the data as a whole are showing.

Categorical and quantitative data

Before choosing or interpreting a graph, determine what type of data you have.

Categorical

Places individuals or objects into groups. The categories describe qualities, not measurements.

favorite subjecttype of vehicleeye colorpayment method

Usually compared with a bar graph.

Quantitative

Numerical measurements or counts.

heightagebooks readtest scoretemperature

Can be analyzed with mean, median, range, and standard deviation.

The distinction matters because different graphs suit different types of data.

Choosing an appropriate graph

Bar graph — compare categories

42 Sports 28 Music 16 Drama 24 Art

Which extracurricular activity students prefer. Categories, not numerical intervals — so a histogram wouldn’t be appropriate. Bars have gaps.

Histogram — distribution of numerical data

3 50–59 7 60–69 12 70–79 15 80–89 8 90–99

Test scores grouped into intervals. The bars touch because the intervals form a continuous scale.

Box plot — center and spread

10 18 24 30 42

Summarizes a dataset with five values: minimum, Q1, median, Q3, maximum. Especially useful for comparing two or more datasets.

Scatterplot — relationship between two variables

Each point is one observation, such as hours studied vs. test score. Shows trends and relationships.

Reading tables carefully

Time Customers
9–11 a.m. 35
11 a.m.–1 p.m. 52
1–3 p.m. 48
3–5 p.m. 65

Which period had the most customers?

3–5 p.m., with 65.

How many visited during the first two periods combined?

35 + 52 = 87

When reading a table

Pay attention to both the row or column labels and the units.

Interpreting histograms

5 10–14 12 15–19 18 20–24 9 25–29 4 30–34

Ages of participants in a program.

Which interval has the most participants?

20–24, with a frequency of 18.

How many participants in total?

5 + 12 + 18 + 9 + 4 = 48

⚠ What a histogram doesn’t tell you

Values are grouped, so you usually can’t determine individual data values. If 18 people are in the 20–24 group, you don’t know how many are 20, 21, 22, 23, or 24. Don’t claim more precision than the graph provides.

Interpreting box plots

10 18 24 30 42

Min10Q118Median24Q330Max42

Range = 42 − 10 = 32IQR = 30 − 18 = 12

The box is the middle 50% of the data, so about half the values fall between 18 and 30.

Comparing box plots

60 72 80 86 96 Class A 55 65 80 92 98 Class B

Both classes have the same median, 80. But the boxes differ in width:

Class A: IQR = 86 − 72 = 14Class B: IQR = 92 − 65 = 27

The middle half of Class B’s scores is more spread out. A higher median means a higher center; a wider box means greater spread in the middle 50%.

Scatterplots and trends

Positive association

Points rise from left to right. As hours studied increase, scores tend to increase — an overall trend, not a guarantee for every student.

Negative association

Points fall from left to right. As a car’s age increases, its resale value tends to decrease.

No clear association

No obvious upward or downward pattern.

Strength of an association

Stronger

Points lie close to a clear line.

Weaker

Same direction, but the points are widely scattered.

Lines of best fit

Line of best fit — a line that approximates the overall trend in a scatterplot. It can be used to estimate values.
ProblemA line of best fit relating hours studied (x) to test score (y) is y = 5x + 60. Predict the score for a student who studies 4 hours.
y = 5(4) + 60substitute y = 80a prediction, not an exact value — the trend is approximate

Outliers

An outlier lies noticeably far from the rest of the data — here, one point well below an otherwise clear upward trend.

Outliers can affect:

the meanthe rangestandard deviationthe apparent trenda line of best fit

Don’t assume an outlier is a mistake. It may be a legitimate unusual observation.

Association does not prove causation

A scatterplot can show that two variables are associated, but not that one causes the other. If people who exercise more tend to report higher energy, that’s an association — the graph alone doesn’t prove exercise is the reason. Other factors could be involved.

The principle

Association does not necessarily imply causation. Avoid conclusions that go beyond the evidence shown.

TSI strategy: read the labels before the data

1What does each axis, row, or column represent?
2What units are being used?
3What does each bar, point, interval, or section represent?

Then read the question carefully.

Asks for a trend

Look at the overall pattern, not one point.

Asks for an exact value

Locate the right point, bar, row, or column.

Asks you to compare datasets

Consider both center and spread.

This prevents one of the easiest data-analysis mistakes: performing the right calculation on the wrong values.

🔑 Key tip: match the graph to the question

Bar graph

compare categories

Histogram

distribution of numerical data

Box plot

center and spread

Scatterplot

relationship between two variables

Don’t just read individual values. Ask what the display tells you about the overall dataset.

Data Interpretation Review Quiz