AMERICAN MILITARY UNIVERSITY • AMU • MATH120
MATH120 Guide: Turn a Statistical Question into a Data and Variable Map
Start by naming the observational unit, target population, observed sample, and every variable needed by the question. Classify a variable from what it means: categorical values name groups, while quantitative values record counts or measurements for which arithmetic differences are meaningful. Then identify each variable’s role, unit or category set, missing-value rule, and connection to the requested comparison. This map prevents a polished summary of the wrong data from masquerading as an answer.
Decision resource
Question-to-Data Map
A structured inventory that connects the evidence question to observational unit, population, observed sample, variable meaning, data type, role, and quality checks.
Step 1
- map field
- Question
- decision question
- What group, variable, relationship, and timeframe must the result address?
- evidence to record
- A learner-written evidence question
- warning signal
- The proposed statistic does not answer a named part
Step 2
- map field
- Observational unit
- decision question
- What does one case or row represent?
- evidence to record
- The unit and any repeated-observation structure
- warning signal
- Rows are counted as independent units without support
Step 3
- map field
- Population and sample
- decision question
- Who is targeted, reachable, and actually observed?
- evidence to record
- Source, eligibility, selection, and response notes
- warning signal
- The conclusion names a wider group than the source supports
Step 4
- map field
- Variable meaning
- decision question
- Does the value name a category or measure an amount?
- evidence to record
- Definition, unit or category set, role, and missing rule
- warning signal
- A numeric identifier is averaged or an ambiguous code is used
Step 5
- map field
- Alignment
- decision question
- Can every word in the conclusion trace to observed evidence?
- evidence to record
- Claim-to-field cross-check
- warning signal
- An unsupported population, comparison, or timeframe appears
Write a statistical question that expects variation
A statistical question anticipates that observations may differ. It identifies a group or process and asks about a distribution, comparison, association, or uncertain outcome rather than one fixed fact. Rewrite the goal so a reader can tell who or what is being studied, what is recorded, and which feature matters. “What is the value?” is not enough when the evidence consists of many cases. Ask what the values reveal about typical behavior, differences, proportions, or relationships.
Keep the scope visible. A question about customers at one fictional location during one month cannot silently become a claim about every customer everywhere. When a comparison is intended, name the groups and common outcome. When an association is intended, name both variables without assuming one causes the other. This preparation is original guidance, not an official classroom template.
Separate observational unit from variable
The observational unit is the individual person, object, event, place, or time period represented by one case. A variable is an attribute recorded for each unit. Confusing the two leads to statements such as calling “age” the sample or treating “patients” as a variable. Write a simple sentence: one row represents ___; for each row, the dataset records ___.
Datasets can have nested or repeated observations. Several readings from the same device are not automatically independent devices, and several purchases by one customer are not automatically separate customers for every question. Before counting rows as sample size, ask which unit the conclusion concerns and whether repeated cases require a different interpretation.
Name the population, available source, and observed sample
The target population is the wider group the question concerns. The observed sample is the set of units actually recorded. Between them may sit a sampling frame, administrative database, eligibility rule, response process, or convenience source. Those layers determine whose experiences can be represented. A learner should not write “the population was sampled” without explaining the connection.
Create three labels: target, reachable, observed. Target names the intended group; reachable names who could enter through the available source; observed names who actually supplied usable values. Differences among the three are not merely procedural. They become limitations on generalization and may reveal coverage or nonresponse concerns that no later calculation can erase.
Classify by meaning, not storage format
Categorical variables place units into labels. Some categories have no natural order; others have a meaningful order but unknown spacing. Quantitative variables record counts or measurements whose numerical differences carry meaning. A numeric code for region, account, or response category remains categorical if arithmetic on the code is meaningless. Conversely, a count stored as text remains quantitative in concept.
Use the arithmetic test carefully. Would subtracting two values describe a meaningful difference in the attribute? Would an average have an interpretable unit? If not, the values may be identifiers or category codes. For ordered categories, the order can support rankings or cumulative summaries, but treating adjacent labels as equally spaced needs evidence rather than convenience.
Identify outcome, explanatory, grouping, and context roles
Variable type and variable role answer different questions. A quantitative measurement can be an outcome in one analysis and an explanatory variable in another. A categorical variable may define comparison groups, mark a subgroup, or act as the outcome whose proportions are compared. State the role relative to the current question instead of assigning a permanent label.
Also record context variables that may affect interpretation, such as timeframe, location, or measurement conditions. Their presence does not prove a causal mechanism, and their absence may limit a comparison. The goal at this stage is an inventory: which variables directly answer the question, which organize the comparison, and which describe conditions that a careful reader needs.
Inspect missingness, coding, units, and measurement definition
A correct type label does not guarantee usable data. Check category definitions for overlap, verify that units are consistent, distinguish zero from missing, and identify impossible or ambiguous values. A missing entry may mean “not recorded,” “not applicable,” “declined,” or a technical failure. Combining those meanings can change a proportion or distribution.
For quantitative measurements, record precision and unit. For categories, record the complete allowed set and whether multiple selections are possible. For dates or durations, decide which derived quantity the question needs and preserve the original reference. Documenting these choices creates an audit trail and avoids silent cleaning decisions that manufacture a cleaner story.
Connect the map to an appropriate representation plan
Once the variable inventory is stable, sketch what each possible representation would encode. A frequency table needs a categorical variable and a defensible denominator. A bar chart needs distinct category labels and a scale that preserves honest length comparison. A histogram needs quantitative values, units, and justified interval choices. A scatterplot needs paired quantitative observations for the same units. State what each axis, row, mark, or cell represents before opening software.
Reject displays that answer a different question. A chart of record identifiers has no statistical meaning, and a mean of response codes does not become useful because software computed it. When two groups are compared, confirm that both use the same definition, timeframe, eligibility rules, and measurement unit. If the data cannot support a proposed representation, document the missing evidence instead of substituting a visually attractive output.
Finish with a summary contract: name the variable, valid denominator or unit, planned display, proposed center or proportion, accompanying spread or uncertainty, and limitation to carry into the conclusion. This contract keeps later calculations attached to the original question and makes mismatches easier to catch during review. Add the anticipated interpretation in plain language and confirm that no field, category, or unit is being asked to support meaning it was never designed to record.
Fictional example: digits can still be categories
Imagine a fictional transit survey in which travel mode is stored as 1, 2, or 3 and commute time is stored in minutes. The travel-mode codes identify categories; their average would not describe an actual mode. Commute time is quantitative because differences in minutes have meaning. The observational unit is one responding commuter, not one mode or one minute. The target group, reachable survey audience, and actual respondents must still be distinguished before generalizing.
This small illustration is intentionally incomplete. It contains no current AMU prompt, dataset, answer, or submission-ready calculation. A learner would still need to define the question, verify the coding guide, inspect missing responses, and decide which summaries address their own evidence.
Run the Question-to-Data alignment check
Read the question and underline every population, variable, comparison, and timeframe term. Then point to the corresponding field or source note. If a term has no evidence, revise the question or obtain appropriate data; do not fill the gap with assumption. Next, test each proposed summary against the variable meaning. Category counts require clear denominators, while quantitative summaries require meaningful units and distribution checks.
Finally, draft a provisional conclusion shell without numbers: “Among the observed ___, the distribution/proportion/relationship of ___ showed ___ during ___.” If the shell forces a group, variable, or timeframe that was never observed, the alignment problem is visible before the calculation begins. Save the corrected map beside the analysis so later revisions preserve variable definitions and the evidence boundary.
Use the map to support learner-owned analysis
Domyclass can help distinguish a unit from a variable, test whether a code is categorical, or review a learner-created data dictionary. It can ask diagnostic questions about the target population and missing values. It will not reproduce a current graded prompt, classify every variable in a live assessment, select quiz answers, or write the final analysis. The learner builds the map, performs the work, explains the choices, and submits only what they understand.
Related MATH120 resources
Get Help With MATH120 Introduction to Statistics at American Military University
Get targeted MATH120 help and improve your grades.
Get MATH120 HelpSources & updates
- American Military University: MATH120 Introduction to Statistics course schedule
- American Public University: MATH120 Introduction to Statistics course schedule
- American Public University System: Mathematics undergraduate course descriptions
- American Public University System: General Education requirements
Published by Domyclass • Updated August 2026