Psychometrics and test construction
Introduced| English |
|---|
| construct validity/ˈkɒnstrʌkt væˈlɪdɪti/ |
| standard error of measurement/ˈstændəd ˈerə ɒv ˈmeʒəmənt/ |
A decision before an answer
- A scale can be perfectly consistent and still measure the wrong thing.
- Your goal: Distinguish reliability types by the design that produced them.
Name reliability by design
- Reliability is consistency, and its type is named by the design that produced it: the same test twice gives test-retest reliability; two similar forms give alternate-forms reliability; consistency within one sitting (split-half, Cronbach alpha) is internal consistency; agreement between scorers is inter-rater reliability.
- The Spearman-Brown formula r_full=2r/(1+r) rescales a half-test correlation to full test length: r=.80 becomes about .89.
Odd and even halves of a test correlate r=.80. After Spearman-Brown correction, the estimated full-test reliability is about:
r_full=2r/(1+r)=1.6/1.8≈.89.
Separate the validity types
- Validity is evidence for the proposed interpretation, not of the test itself: content validity asks whether items cover the domain; criterion validity relates scores to an outcome, concurrently or predictively; construct validity assembles convergent and discriminant evidence that the score tracks the intended trait.
- Reliability caps validity but never establishes it: a consistent scale can systematically measure something else.
Evidence that a score tracks the intended trait rather than a different one is primarily:
Construct validity rests on convergent and discriminant evidence; the others are reliability or sampling properties.
Standardise the scores
- Standard scores re-express raw scores against a norm group: z=(X−mean)/SD, T=50+10z, and Wechsler deviation IQ=100+15z. A z of +2 gives IQ 130 and about the 98th percentile.
- The standard error of measurement SEM=SD×sqrt(1−r) converts reliability into a score band: SD=15 with r=.91 gives SEM=15×sqrt(.09)=4.5.
Two similar vocabulary tests given a week apart with r=.90 show alternate-forms reliability, not test-retest, because the form changed. With SD=15 and reliability .91, SEM=4.5 points. A z of +2 corresponds to IQ 130. A split-half correlation of .80 corrects to 2×.80/1.80≈.89 at full length.
A Wechsler score of 130 corresponds to z=____.
130=100+15z gives z=2.
Analyse the items
- Item analysis checks each item: difficulty p is the proportion answering correctly, and discrimination is the item-total correlation.
- Items that nearly everyone or no one answers correctly carry little information; negative discrimination flags a likely miskeyed or misleading item.
Calling a two-different-forms correlation test-retest reliability, or citing high reliability as if it were validity evidence.
Which answer fits this case?
Distinguish reliability types by the design that produced them
A highly reliable test must be valid for its stated purpose.
Consistency does not establish that the intended interpretation is correct.
Keep the distinctions
- standard error of measurement 测量标准误 — The expected spread of observed scores around the true score, SD×sqrt(1−r).
- construct validity 构想效度 — Convergent and discriminant evidence that a score reflects the intended trait.
- Distinguish reliability types by the design that produced them.
- Separate content, criterion and construct validity evidence.
- Compute standardised scores and the standard error of measurement.
Match each term with its precise meaning in this lesson.
Keep the distinctions stated in the teaching example.
Put this lesson’s reasoning or event sequence in order.
The order follows the stated process; check each stage before the next.