Concise answer
The measurement domain is about entitlement: given this design and this result, what may be said? Designs differ in what they license — description, prediction, or causation — and the difference is created by two independent randomisations that students routinely merge. Statistics then quantify a result without ever converting it into a stronger claim than the design supports, and the classical psychometric distinction between reliability and validity does the same work at the level of the instrument. Intelligence testing is the domain's worked historical example, and research ethics is the set of constraints under which all of it operates.
Definitions
- Random sampling versus random assignment
- Random sampling draws participants from a population and supports generalisation. Random assignment allocates participants to conditions and supports causal inference. A study can have either, both, or neither.
- Operational definition
- A statement of the exact procedure by which a construct is measured or manipulated in this study. It is what makes replication and criticism possible.
- Confounding variable
- A variable that differs systematically between conditions alongside the independent variable, so that the effect cannot be attributed to the manipulation.
- Internal and external validity
- Internal validity is confidence that the manipulation caused the change; external validity is confidence that the finding extends beyond this sample and setting. Tight control usually raises the first and can lower the second.
- Reliability
- Consistency of measurement — across occasions (test–retest), across raters (inter-rater), or across items (internal consistency). A measure can be perfectly consistent and still measure the wrong thing.
- Validity
- Whether the instrument measures the construct it claims to. Content, criterion (concurrent and predictive), and construct validity are the standard forms; face validity is only about appearance.
- Type I and Type II error
- A Type I error rejects a true null hypothesis — a false positive, with probability set by alpha. A Type II error fails to reject a false null — a false negative, with probability beta, so power equals one minus beta.
- Effect size
- A measure of how large a difference or relationship is, independent of sample size. Statistical significance says a result is unlikely under the null; effect size says whether it matters.
Intuition
Two randomisations do two different jobs and neither substitutes for the other. Randomly selecting who takes part tells you the sample resembles the population, which is a claim about generalisation. Randomly allocating those participants to conditions tells you the groups did not differ beforehand, which is a claim about cause. A tightly controlled experiment on one university's undergraduates has the second without the first; a national survey has the first without the second. Almost every methodology item is built on that split.
Statistical significance is a statement about a hypothetical world, not about this one. A p-value answers: if there were truly no effect, how often would data this extreme arise? A small value makes the no-effect account uncomfortable. It does not report the probability that the null is true, it does not report the probability of replication, and because it depends on sample size, it does not report how large the effect is. Reading every 'p < .05' as 'the no-effect story fits badly' keeps all three misreadings out.
Reliability and validity stand in an asymmetric relationship worth stating out loud. A scale that reads three kilograms heavy is perfectly reliable and completely invalid. But a scale that returns a different number every time cannot be measuring anything consistently, so it cannot be valid either. Reliability is necessary for validity and nowhere near sufficient — which is why an item offering high reliability as evidence of validity is offering the wrong direction of the implication.
Concept walkthrough
Designs form a ladder of entitlement. A case study describes one person or setting in depth and is invaluable for rare phenomena and hypothesis generation, but supports no general conclusion. Naturalistic observation records behaviour where it occurs and buys realism at the cost of control, with observer effects and observer bias as its characteristic threats. Surveys reach large samples cheaply and measure self-report, which is not the same as behaviour and is subject to social desirability. Correlational research measures two variables as they occur and supports prediction, but not causation, because of the directionality problem and the possibility of a third variable. Only an experiment — a manipulated independent variable, a controlled comparison condition, and random assignment — supports the conclusion that the manipulation produced the change. Quasi-experiments compare pre-existing groups such as sexes or diagnoses; because assignment was not random, they inherit the correlational limits no matter how experimental they look.
Within an experiment, the threats have names, and naming them is the skill. A confound is any variable that varies systematically with the independent variable. Demand characteristics are cues that let participants infer the hypothesis and adjust. Experimenter expectancy is the researcher's unintended influence on participants or on scoring, which is why double-blind procedures exist. Placebo effects require a placebo control to isolate. Order effects in within-subjects designs are handled by counterbalancing, and selection bias in between-subjects designs is what random assignment prevents. Between-subjects and within-subjects designs trade off against each other: within-subjects removes individual differences and needs fewer participants but risks order effects; between-subjects avoids order effects but requires more participants and relies on randomisation to equate the groups.
Descriptive statistics summarise a sample. The mean is the balance point and is pulled towards extreme values; the median is the middle score and is not; the mode is the most frequent. In a positively skewed distribution — a long right tail, such as income — the mean sits above the median, and in a negatively skewed distribution it sits below, which is why the median is reported for skewed data. Variability is summarised by the range, the interquartile range, and above all the standard deviation, the typical distance of a score from the mean. In a normal distribution roughly 68 per cent of scores fall within one standard deviation of the mean, 95 per cent within two, and 99.7 per cent within three, so a z-score — the number of standard deviations a score sits from the mean — converts any score into a position in that distribution.
Correlation is the domain's most abused statistic. The coefficient runs from −1 to +1, the sign gives direction and the absolute value gives strength, so −0.8 describes a stronger relationship than +0.5. Squaring it gives the proportion of variance shared, which is why an apparently respectable r of 0.30 accounts for only nine per cent of the variance. A correlation coefficient describes a linear relationship, so a strong curvilinear relationship can produce a coefficient near zero. Restriction of range attenuates correlations, and a single extreme outlier can create or destroy one. And it never licenses a causal claim: the relationship may run the other way, or both variables may depend on a third.
Inferential statistics ask whether a result would be surprising if nothing were going on. The null hypothesis states there is no effect; the p-value gives the probability of data at least this extreme under that assumption; and when it falls below the pre-set alpha, conventionally .05, the null is rejected. Two errors follow: rejecting a true null is a Type I error, whose long-run rate is alpha, and failing to reject a false null is a Type II error, with rate beta. Power, the probability of detecting an effect that is really there, is one minus beta, and it rises with sample size, with effect size, with a less stringent alpha, and with measurement precision. Lowering alpha therefore reduces Type I errors while increasing Type II errors — the trade-off is the point. Test selection is mechanical once the design is clear: a t-test compares two means, one-way analysis of variance compares three or more, chi-square handles frequencies in categories, and correlation or regression handles two measured continuous variables. Effect size measures such as Cohen's d, with the conventional landmarks near 0.2, 0.5, and 0.8, answer the question significance cannot.
Psychometrics applies the same rigour to instruments. Reliability comes in three flavours — stability across time, agreement across raters, and internal consistency across items — and internal consistency is usually reported as coefficient alpha. Validity asks whether the instrument measures the intended construct: content validity covers the domain adequately, criterion validity predicts a relevant outcome measured now (concurrent) or later (predictive), and construct validity assembles the whole pattern of evidence. Face validity, whether the test merely looks right, is the weakest form and is sometimes deliberately sacrificed. Standardisation means uniform administration and scoring plus norms from a representative sample, without which an individual score cannot be interpreted at all. The standard error of measurement expresses the uncertainty around an individual score and is why scores are best reported as bands.
Intelligence is the domain's extended case study. Spearman's factor analysis suggested a general factor underlying performance across tasks alongside task-specific factors; Thurstone argued for several primary mental abilities instead; Cattell distinguished fluid reasoning from crystallised knowledge; Gardner proposed multiple relatively independent intelligences; and Sternberg's triarchic account separated analytic, creative, and practical intelligence. Measurement began with Binet's practical task of identifying children needing additional schooling, produced the mental-age idea, and was converted into the ratio intelligence quotient of mental age over chronological age times one hundred — a formula that fails for adults, since chronological age keeps rising while performance does not. Modern tests therefore use a deviation score, placing an individual against the distribution of scores for their own age group, with the Wechsler scales set to a mean of 100 and a standard deviation of 15. Two further facts recur: the Flynn effect, the rise in raw test performance across generations that forces periodic renorming, and the correct reading of heritability as a statement about the proportion of variance within a population attributable to genetic differences, not about how much of an individual's ability is inherited.
Ethics constrains all of it. Human research requires informed consent given by a competent participant who understands what is involved, voluntary participation with the right to withdraw at any time without penalty, protection from harm beyond that of ordinary life, confidentiality of data, and debriefing afterwards. Deception is permitted only when the study cannot otherwise be done, when the risk is justified, and when full debriefing follows. Institutional review boards exist to make that judgement independently of the researcher, and animal research is governed by its own committees and standards of care. The historical cases that drove these requirements — medical research conducted without consent, and obedience and prison studies whose participants experienced real distress — are examinable both as history and as illustrations of which specific principle was breached.
After this page, you should be able to
- State what each research design licenses and refuse the causal option when the design cannot support it.
- Distinguish random sampling from random assignment and say which conclusion each one underwrites.
- Compute and interpret a z-score, and use the normal distribution's landmark percentages to place a score.
- Choose the appropriate inferential test from the number of groups and the type of data.
- State what a p-value does and does not mean, and separate significance from effect size and from power.
- Distinguish reliability from validity, and explain why a reliable measure may still be invalid while an unreliable one cannot be valid.
- Compare the major theories of intelligence and interpret an IQ score as a position in a distribution rather than as a quantity.
- List the ethical requirements for human research and identify which one a described protocol violates.
Formulas and assumptions
z-score and the normal distribution
Variables
- X: the individual score
- M: the mean of the distribution
- SD: the standard deviation
Assumptions
- The landmark percentages hold only for an approximately normal distribution.
- A z-score is a position, not a quantity: z = 2 means two standard deviations above the mean, whatever the raw units are.
Correlation and shared variance
Variables
- r: the linear correlation coefficient
- r squared: the coefficient of determination
Assumptions
- r describes a linear relationship, so a strong curvilinear pattern can give r near zero.
- A correlation never establishes causation: the direction may be reversed or a third variable may drive both.
Two randomisations, two entitlements
random sampling => generalisation to the population; random assignment => causal inference within the study
Variables
- sampling: how participants were drawn from the population
- assignment: how participants were allocated to conditions
Assumptions
- The two are independent: a study can have one without the other, and each supports only its own conclusion.
- Quasi-experiments compare pre-existing groups, so they lack random assignment however experimental the procedure looks.
Error and power
Type I = reject a true null (rate alpha); Type II = fail to reject a false null (rate beta); power = 1 - beta; lowering alpha reduces Type I and increases Type II
Variables
- alpha: the significance level chosen in advance, conventionally .05
- power: the probability of detecting a real effect
Assumptions
- Power rises with sample size, with effect size, with a less stringent alpha, and with measurement precision.
- Failing to reject the null is not evidence that the null is true; it may only mean the study lacked power.
Inferential test selection
two means => t-test; three or more means => ANOVA; frequencies in categories => chi-square; two measured continuous variables => correlation or regression
Variables
- means: continuous outcome compared across groups
- frequencies: counts of cases falling into categories
Assumptions
- Paired or repeated measurements call for the within-subjects version of the test rather than the independent-groups version.
- The test follows from the design; choosing a test cannot rescue a design that never licensed the conclusion.
Reliability and validity relation
reliability = consistency; validity = measuring the intended construct; reliable but invalid is possible, valid but unreliable is not
Variables
- reliability forms: test-retest, inter-rater, internal consistency
- validity forms: content, criterion (concurrent, predictive), construct
Assumptions
- Reliability is necessary but far from sufficient for validity — a consistently biased instrument is perfectly reliable.
- Face validity concerns appearance only and carries no psychometric weight.
IQ: ratio versus deviation
Variables
- MA: performance expressed as the age at which it is typical
- CA: the person's actual age
Assumptions
- The ratio formula breaks down in adulthood because chronological age keeps rising while test performance does not.
- A deviation score is meaningful only relative to the norm group it was computed against.
Worked example
Reading a significant result for exactly what it licenses
Ninety undergraduates from one university volunteer for a study on study technique. They are randomly assigned to a retrieval-practice group or a rereading group. The retrieval group's mean test score is 78 and the rereading group's is 72; the pooled standard deviation is 12 and the difference is significant with p = .04. The authors conclude that retrieval practice causes better learning in students generally. Evaluate the conclusion, compute the effect size, and state what p = .04 does and does not mean.
- 1Separate the two randomisations before anything else. Participants were randomly assigned to conditions, but they were not randomly sampled from any wider population — they are volunteers from a single university.
- 2Take the causal half of the claim. Random assignment means the two groups should not differ systematically in prior ability, motivation, or anything else, so the difference in outcome is attributable to the manipulated variable. The causal conclusion is supported within this study.
- 3Take the generalisation half. Nothing licenses extending the result to students in general: volunteers from one institution are a convenience sample, and volunteering is itself correlated with motivation. The correct scope is 'in this population, under these conditions'.
- 4Compute the effect size. Cohen's d is the difference in means divided by the pooled standard deviation: (78 − 72) / 12 = 6 / 12 = 0.50, a medium effect by the conventional landmarks.
- 5Interpret the effect size alongside the significance. A difference of half a standard deviation is worth having, and stating it protects against both errors an item can offer — dismissing a significant result as trivial, and inflating it into a transformation.
- 6State what p = .04 means. If retrieval practice and rereading truly produced no difference, data at least this extreme would arise about 4 per cent of the time. That is a statement about the frequency of such data under the null hypothesis.
- 7State what it does not mean. It is not the probability that the null hypothesis is true; it is not the probability that the finding would replicate; and it is not a measure of how large the difference is, since with ninety participants a much smaller difference could also have reached significance.
- 8Say what would strengthen each half of the claim separately. For the causal claim: replication with the same design and a pre-registered analysis. For the generalisation: sampling across institutions and student populations. The two weaknesses have different remedies, and confusing them is what the wrong options are built on.
The causal half is supported by random assignment; the generalisation to students in general is not, because the sample was self-selected from one university. Cohen's d = (78 − 72) / 12 = 0.50, a medium effect. And p = .04 means data this extreme would occur about 4 per cent of the time if there were truly no difference — not that the null is 4 per cent likely, nor that the result would replicate 96 per cent of the time.
Common traps
- Merging random sampling with random assignment. The first licenses generalisation, the second licenses causal inference, and a study may have either without the other.
- Reading a p-value as the probability that the null hypothesis is true. It is the probability of data this extreme assuming the null is true.
- Treating statistical significance as evidence of a large effect. With a large sample a trivial difference can be significant, which is exactly why effect size is reported.
- Treating a non-significant result as proof of no effect. Failure to reject may reflect low power rather than absence of an effect.
- Assuming lowering alpha simply makes the study better. It reduces Type I errors and increases Type II errors at the same time.
- Reversing the skew rule. In a positively skewed distribution the mean lies above the median, which is why the median is preferred for skewed data.
- Comparing correlations by sign. An r of −0.8 is stronger than an r of +0.5; only the absolute value carries strength.
- Confusing r with r squared. An r of 0.30 corresponds to nine per cent of shared variance, not thirty.
- Concluding that a near-zero correlation means no relationship. A strong curvilinear relationship produces a coefficient near zero.
- Taking high reliability as evidence of validity. A consistently biased instrument is perfectly reliable and entirely invalid.
- Applying the ratio IQ formula to adults. Modern scores are deviation scores relative to an age-group norm, not a ratio of mental to chronological age.
- Reading heritability as the share of an individual's ability that is inherited. It describes the proportion of variance within a population under the conditions studied.
- Treating a quasi-experiment as an experiment. Comparing pre-existing groups leaves the same third-variable problem correlational research has.
Related pages and practice
Question depth and domain coverage vary by exam. Practice answers are checked after submission.
Sources
- GRE Subject Test Content and Structure — ETS. Accessed 2026-07-06. Use as a cited source for exam facts; do not imply affiliation or reproduce protected test material.
- Psychology 2e, Section 2.2: Approaches to Research — OpenStax. Accessed 2026-08-03. OpenStax textbook content is CC BY-NC-SA 4.0; attribute and avoid verbatim reuse beyond short cited references.
- Psychology 2e, Section 2.3: Analyzing Findings — OpenStax. Accessed 2026-08-03. OpenStax textbook content is CC BY-NC-SA 4.0; attribute and avoid verbatim reuse beyond short cited references.
- Psychology 2e, Section 7.5: Measures of Intelligence — OpenStax. Accessed 2026-08-15. OpenStax textbook content is CC BY-NC-SA 4.0; attribute and avoid verbatim reuse beyond short cited references.
- Psychology 2e, Section 1.2: History of Psychology — OpenStax. Accessed 2026-08-15. OpenStax textbook content is CC BY-NC-SA 4.0; attribute and avoid verbatim reuse beyond short cited references.
Sources and corrections
Sources last checked 2026-08-15Every source cited on this page was checked on the date shown, and we update the page when a source changes. If something looks wrong, tell us and we'll recheck it.