Unit 2 · 2.8 · hard topic
Intelligence and Achievement
Unit 2 is 15–25% of the multiple-choice section. This is a hard topic: intelligence is a debated construct, a score is not a person, and this course does not rank groups. · about 5 minutes with the essay beats.
You do not remember what happened. You remember what you encoded — and what later got in the way.
Hard — this one trips people who only memorized the word
Read the scene first. The term will wait. When it lands, you will be able to use it on a stranger’s story — that is the exam, and that is also why this subject is interesting.
A scene you already lived
Do not hunt the term yet. Let this sit. The exam will hand you a stranger’s version of the same night.
A cousin solves a viral puzzle in twelve seconds. The group chat types genius. A week later a school report returns a number, and someone treats the number as a soul, or as a ranking of families, or as destiny. Intelligence, in this course, is a construct the field debates: a statistical factor some researchers call g, a popular multiple-intelligences story with measurement limits, a triarchic framing of analytical, creative, and practical. An IQ score is a score on a test given under rules. Reliability, validity, standardization, and bias are methods ideas. They tell you what a number can claim. They do not tell you who to admire, and they are not a license to rank ethnic groups.
What the paper actually asks
These are the scoring targets. If you can do them on a new story, you are ready — flashcards can wait.
- Define intelligence as the field debates it — not as a mystical IQ aura.
- Apply reliability, validity, and bias as research-methods ideas.
- Never rank groups. Never treat a score as a person.
The mechanism
Once you can walk this, you will start seeing it at dinner, in a group chat, in your own delay. That is the fun of this course.
The field does not all mean the same thing by intelligence. Spearman’s g is a statistical-factor idea: scores on many cognitive tasks tend to correlate, and g is a name for that shared variance — a pattern in numbers, not a fluid in the blood and not a soul. Gardner’s multiple-intelligences framing (linguistic, musical, bodily-kinesthetic, and others) is a popular alternative. Treat it as a map with a measurement limit: it is hard to show these as separate intelligences rather than as talents, skills, or interests, and a viral ‘which intelligence are you’ quiz is not a test. Sternberg’s triarchic framing — analytical, creative, practical — is another map, not a brain scan and not a third IQ number you should memorize as destiny.
IQ is a score on a particular standardized test, compared with a reference sample, not a destiny and not a person. Achievement tests claim to measure learned knowledge or skill in a domain (what this course already taught). Aptitude tests claim to predict later learning or performance. Those are different claims; a high score on one is not automatic proof of the other.
Methods language is the exam’s real lever. Reliability: consistency — similar scores across time, items, or forms if the construct is stable. Validity: whether the test measures the construct it claims (construct), covers the relevant content (content), and predicts what it should (predictive). A test can be reliable (the number is stable) and still weakly valid for a particular use. Standardization: same procedures, scores interpreted against a norming sample. Bias, as a research-methods idea: a test may not measure the same construct across groups or contexts, or it may predict differently — that is a problem for the instrument and the use, not a prompt to rank ethnic groups. This course does not rank ethnic groups. This course does not treat a score as a person.
Construct
what the field is even arguing
Reliability
consistency of the measure
Validity
does it measure the claim?
Standardization
same procedure, compared norms
Bias (methods)
not the same construct across groups
Achievement vs aptitude
learned vs predicted — different claim
Words worth owning
A definition is a tool. The confusable is the distractor. Learn both and the clever wrong answer stops working.
g (general intelligence as a factor)
A statistical-factor idea: shared variance among many cognitive tests, named as general intelligence by some researchers.
Confusable. Not a fluid, not a soul, not proof of destiny. Not the same as a viral puzzle score.
Multiple intelligences
A popular alternative framing that names several intelligences (for example linguistic, musical, bodily-kinesthetic). Measurement of them as separate intelligences is the limit.
Confusable. Not an official sorting hat. A talent or interest is not automatically a separately measured intelligence.
Triarchic framing
Sternberg’s map of analytical, creative, and practical aspects of intelligence — another framing, not a second IQ mystique.
Confusable. Not three brain lobes. Not a license to skip reliability and validity.
IQ score
A number from a standardized intelligence test, interpreted against a reference sample. A score, not a person and not a destiny.
Confusable. Not achievement in a course. Not proof of genius from one viral puzzle.
Reliability
Consistency of scores — across time, items, or forms — if the construct is stable.
Confusable. Reliability is not validity. A consistent score can still fail to measure the intended construct.
Validity
How well a test measures the construct it claims (construct), covers relevant content (content), and predicts a later criterion (predictive).
Confusable. Not ‘the number feels true.’ Not automatic just because the test is famous or consistent.
Standardization
Uniform procedures and interpretation of scores against a norming sample.
Confusable. Not bias, and not a claim that the norming sample represents every context. Uniform rules are not the same as fairness.
Bias (as a methods idea) / achievement vs aptitude
Bias: a test may not measure the same construct across groups or contexts, or may predict differently — a research-methods problem. Achievement tests claim learned skill; aptitude tests claim to predict later learning. Different claims.
Confusable. Bias is not a slogan that ‘all tests are destiny,’ and it is not a prompt to rank ethnic groups. Achievement is not aptitude.
A study you can actually use
Classic or labeled hypothetical. Either way: what was done, what was found, what it cannot claim.
Classic (classroom-common) · early 1900s (classroom-common)
Binet’s practical school-help purpose (classroom-common history)
- What was done
- Classroom-common account: Alfred Binet developed tasks to help identify children who might need extra help in school, not to crown genius or to rank nations.
- Finding
- Early testing was a practical school tool. Later IQ culture often drifted into mystique Binet did not own. A score began as an aid for instruction, not as a soul.
- Limit — what it cannot claim
- A historical origin story is not a modern validity coefficient. It does not license ranking groups, inventing national averages, or treating today’s tests as destiny. Classroom history is simplified; do not invent a journal or a mean.
- How AAQ would probe it
- Identify the original claimed use (help in school). Limit: origin is not modern construct validity. Application: a score is a tool with a use, not a person.
Labeled hypothetical
Hypothetical two-form reliability and predictive coefficients (original practice source, not a published paper)
Original practice source, not a published paper.
- What was done
- A school tries two parallel forms of a reasoning test, two weeks apart, with the same 80 students, and later records end-of-year course marks. Form A mean = 101; Form B mean = 100. Correlation of Form A with Form B = 0.87. Correlation of Form A with later course marks = 0.19. A second table compares morning and afternoon sittings of Form A (same students, counterbalanced): morning mean = 102, afternoon mean = 99.
- Finding
- High form-to-form agreement supports reliability for this sample. The weak correlation with course marks is weak evidence of predictive validity for those marks. A three-point sitting-time shift is a context note, not a ranking of people. Nothing here ranks ethnic groups, and nothing here proves genius.
- Limit — what it cannot claim
- Labeled hypothetical. Numbers exist only in this practice card. 80 students in one school. Course marks are one criterion, not ‘intelligence.’ Do not treat 101 as a soul. Do not invent national averages. Do not cite as a real paper.
- How AAQ would probe it
- Practice 3: say what each number allows. 0.87 → consistency (reliability). 0.19 → weak prediction of that criterion. Means near each other → forms are similar on average, not a contest of worth. Bias as methods: if a form behaved differently across contexts (morning vs afternoon), that is a use-and-context problem — still not a group ranking.
Someone’s Tuesday
If it only lives in the textbook, it will not survive a novel stem. This is the idea wearing ordinary clothes.
A cousin’s viral puzzle is entertainment. It is not an IQ test, not a reliability study, and not a ranking of your family. A music exam measures achievement in that syllabus; it does not, by itself, measure g, and it does not prove a ‘musical intelligence’ as a separately validated construct unless someone actually validates it. A standardized entrance paper given with the same instructions and a published norm is a standardized procedure — still not destiny. If a test written in one school’s dialect is used in another context, bias as a methods idea is on the table: does it measure the same construct here? That question is allowed. Ranking ethnic groups is not.
Where clever students go wrong
The trap is usually a neighboring term that almost fits. Name it so it stops feeling smart.
‘My cousin is a genius because of one viral puzzle’ — no. Genius is not a prize this page awards, and a puzzle is not a standardized test. A score is not a person. This course does not rank ethnic groups and does not invent national averages. Reliability is not validity. Multiple-intelligences quizzes on social media are not measurement. IQ mystique — treating a number as a soul, a brand, or a forecast of who deserves care — is out of bounds. If a stem invites you to rank groups, you are being tested on whether you refuse.
A stem, then the move
Watch me work one. Then you will do four without looking back.
Scenario
A school reports Form A mean 101, Form B mean 100, A–B correlation 0.87, and Form A with later course marks 0.19. A comment thread says the test ‘proves who is actually smart’ and starts comparing families. What do the numbers allow, and what must you refuse?
Walkthrough
0.87 is evidence of reliability (forms agree). 0.19 is weak predictive validity for course marks — the test is not doing much of that job here. Means of 101 and 100 are similar averages, not a ranking of souls. Achievement in the course is a different claim from this reasoning score (aptitude-style). Refuse: ranking families or ethnic groups, IQ mystique, treating 101 as a person. Binet’s original school-help story does not rescue the comment thread.
The move
Reliable here; weak as a predictor of those marks. A score is not a person. Do not rank groups. Do not call a viral puzzle or a family comparison science.
Try tonight · 5 minutes
A living experiment, not a vibe
Do this on your actual Tuesday. The paper will later hand you a stranger’s version of the same five minutes.
Find one number people treat as a person (a mock score, a ranking, a puzzle time). Write what reliability would mean for that number, what validity would mean, and one claim you refuse (destiny, group ranking, genius from a clip).
Prove it
Four items. No looking back.
One at a time. After each submit I will show the key, why it is right, and why the others felt clever.
Item 1 of 4 · Practice 1 · Concept application
A cousin solves a viral puzzle in twelve seconds. The group chat types genius and starts ranking families. Which exam-ready reply is best?
Keys A–D select. Enter checks.
Write it · AAQ
Say what the coefficient allows — not who is better
Say the move in your own mouth. If you can write it, you own it — rereading is not the same thing.
Original practice source, not a published paper. Form A–Form B correlation = 0.87. Form A with later course marks = 0.19. Morning mean = 102, afternoon mean = 99, same students. Identify one methods concept the 0.87 is evidence for. Then say what the 0.19 allows you to claim, and state one conclusion about people or groups that these numbers do not allow.
Move: Practice 3. 0.87 → reliability. 0.19 → weak predictive validity for that criterion. Refuse ranking people or groups; refuse treating a mean as a soul.
Where curiosity goes next
Related rooms. Open the one that still tugs.
