19 questions

Assessment and Diagnosis

A test has a standard deviation of 15 and a reliability coefficient of .91. What is its standard error of measurement?

  • a.13.65
  • b.9.0
  • c.1.35
  • d.4.5✓

SEM = SD x the square root of (1 - reliability) = 15 x sqrt(1 - .91) = 15 x sqrt(.09) = 15 x .30 = 4.5. Multiplying 15 by .09 gives 1.35 and 15 by .91 gives 13.65, both of which skip the square root; 9.0 doubles the correct value.

Assessment and Diagnosis

A client obtains a score of 110 on a test whose standard error of measurement is 5. Which range contains the client's true score with about 95% confidence?

  • a.About 108 to 112
  • b.About 100 to 120✓
  • c.About 105 to 115
  • d.About 95 to 125

A 95% confidence band is the obtained score plus or minus 1.96 SEM: 1.96 x 5 = 9.8, so about 100 to 120. Plus or minus one SEM (105 to 115) gives only about 68% confidence, plus or minus three SEMs (95 to 125) gives about 99.7%, and 108 to 112 is narrower than even a single SEM.

Assessment and Diagnosis

Assuming normally distributed scores, a T-score of 60 corresponds most closely to which percentile rank?

  • a.98th
  • b.84th✓
  • c.75th
  • d.60th

T-scores have a mean of 50 and a standard deviation of 10, so 60 is one standard deviation above the mean (z = +1.0). In a normal distribution about 84% of scores fall below z = +1.0. The 98th percentile corresponds to about z = +2.0, and 60th and 75th confuse the T value or a quartile with the percentile.

Assessment and Diagnosis

The correlation between the two halves of a test is .60. Using the Spearman-Brown formula, what is the estimated reliability of the full-length test?

  • a..60
  • b..90
  • c..75✓
  • d..36

For doubling test length, Spearman-Brown gives 2r / (1 + r) = (2 x .60) / (1 + .60) = 1.20 / 1.60 = .75. The half-test correlation of .60 underestimates full-test reliability, .36 is r squared, and .90 overstates the correction.

Assessment and Diagnosis

A screening test has sensitivity of .90 and specificity of .90. It is given to 1,000 people, 10% of whom actually have the condition. What proportion of people who screen positive actually have the condition?

  • a.81%
  • b.50%✓
  • c.10%
  • d.90%

Of 100 true cases, 90 test positive (sensitivity .90). Of 900 non-cases, 10% (90) test positive falsely (specificity .90). Positive predictive value = 90 / (90 + 90) = 50%; recounting, true positives 90 and false positives 90 give the same result. The 90% figure confuses sensitivity with predictive value, and the result shows how a low base rate lowers PPV.

Assessment and Diagnosis

A new anxiety scale correlates .65 with an established anxiety measure and .08 with a vocabulary test. Together these results provide evidence of:

  • a.Test-retest and alternate-forms reliability
  • b.Predictive criterion-related validity
  • c.Content validity
  • d.Convergent and discriminant validity✓

High correlation with a measure of the same construct (convergent) and low correlation with a measure of an unrelated construct (discriminant) are the two sides of construct validity in Campbell and Fiske's (1959) multitrait-multimethod framework, building on Cronbach and Meehl (1955). Reliability concerns consistency of the same test, content validity concerns sampling of the content domain, and predictive validity requires a later criterion.

Assessment and Diagnosis

Scores on a hiring test given at application correlate .40 with supervisor performance ratings collected a year later. This coefficient is evidence of:

  • a.Predictive validity✓
  • b.Internal consistency
  • c.Concurrent validity
  • d.Face validity

When criterion data are collected after the test, the correlation is a predictive validity coefficient, a form of criterion-related validity. Concurrent validity uses criterion data gathered at about the same time as testing, face validity concerns whether a test appears to measure what it claims, and internal consistency is a type of reliability.

Assessment and Diagnosis

A predictor has a validity coefficient of .80 with a criterion whose standard deviation is 10. What is the standard error of estimate?

  • a.3.6
  • b.8
  • c.6✓
  • d.2

Standard error of estimate = SD of the criterion x the square root of (1 - r squared) = 10 x sqrt(1 - .64) = 10 x sqrt(.36) = 10 x .6 = 6. Using 1 - r gives 2, using 1 - r squared without the square root gives 3.6, and 8 multiplies the SD by r.

Assessment and Diagnosis

For a free-response item with no chance of guessing correctly, which item difficulty index (p, the proportion answering correctly) gives the item the greatest potential to discriminate among examinees?

  • a.1.00
  • b..90
  • c..10
  • d..50✓

An item's variance is p(1 - p), which peaks at p = .50 (.50 x .50 = .25), so a middle-difficulty item can separate the most pairs of examinees. Very hard (.10) or very easy (.90) items produce little variance (.09), and an item everyone passes (1.00) has no variance and cannot discriminate at all.

Assessment and Diagnosis

In a three-parameter item response theory model, what does the lower asymptote of an item characteristic curve represent?

  • a.The overall reliability of the full test
  • b.How sharply the item separates higher from lower ability
  • c.The probability of a correct answer at very low ability✓
  • d.The ability level at which the item is of middle difficulty

In the three-parameter logistic model (Lord, 1980), the c parameter, or lower asymptote, is the probability that examinees of very low ability answer correctly, often called pseudo-guessing. The b parameter locates item difficulty on the ability scale, the a parameter reflects the curve's steepness (discrimination), and test reliability is not an item parameter.

Want these explained in order? EPPP Part 1 Study Guide — 2026 Edition — PDF + EPUB, $19.99 · 14-day refund →

Assessment and Diagnosis

On the MMPI-2, a markedly elevated F (Infrequency) scale most suggests:

  • a.Strong ego resources and psychological health
  • b.Subtle defensiveness in a well-educated person
  • c.Overreporting of symptoms or random responding✓
  • d.A naive attempt to appear unusually virtuous

The MMPI-2 F scale is made of items rarely endorsed by the normative sample, so marked elevations suggest overreporting, severe distress, or random or inconsistent responding. A naive attempt to look virtuous is what the L scale detects, subtle defensiveness is associated with K, and elevated F does not indicate psychological health.

Assessment and Diagnosis

Meehl (1954) and the Grove et al. (2000) meta-analysis compared clinical judgment with statistical (mechanical) prediction. Their overall conclusion was that:

  • a.The two methods differed only for personality variables
  • b.Clinical judgment clearly outperformed statistical prediction
  • c.Statistical prediction usually matched or beat clinical judgment✓
  • d.Statistical methods worked only when clinicians lacked interviews

Meehl (1954) reviewed studies showing actuarial prediction as accurate as or more accurate than clinicians, and Grove and colleagues' (2000) meta-analysis of 136 studies found mechanical prediction equal or superior in most studies, with clinical prediction substantially better in only a small minority of studies. Access to interview data did not reverse this advantage.

Assessment and Diagnosis

After two recent, vivid cases of clients who attempted suicide, a clinician begins overestimating how often suicide attempts occur among all new intakes. Which judgment heuristic is most clearly at work?

  • a.Anchoring and adjustment
  • b.Confirmation bias
  • c.Hindsight bias
  • d.Availability✓

Tversky and Kahneman (1974) described the availability heuristic as judging frequency or probability by how easily instances come to mind; recent, vivid cases are highly available. Anchoring involves insufficient adjustment from a starting value, hindsight bias is believing after the fact that an outcome was predictable, and confirmation bias is seeking evidence that fits an existing hypothesis.

Assessment and Diagnosis

Clinicians reported that certain drawing features, such as large eyes, went with suspiciousness, even when those features were paired randomly with symptoms in the case materials they reviewed. This error is called:

  • a.Illusory correlation✓
  • b.The halo effect
  • c.The Barnum effect
  • d.Base-rate neglect

Chapman and Chapman (1967) showed that clinicians and naive observers 'saw' expected relationships between Draw-a-Person features and symptoms that did not exist in the data, which they termed illusory correlation. The Barnum effect is accepting vague personality descriptions as accurate, the halo effect generalizes from one trait, and base-rate neglect ignores prior probabilities.

Assessment and Diagnosis

A licensing exam sets a fixed standard of mastery, and a candidate passes by meeting that standard regardless of how other candidates score. This score interpretation is:

  • a.Percentile-based
  • b.Norm-referenced
  • c.Age-equivalent
  • d.Criterion-referenced✓

Criterion-referenced interpretation compares performance with a defined standard or content domain, whereas norm-referenced interpretation compares a person with a reference group. Percentile ranks locate a person relative to a norm group, and age-equivalent scores relate performance to typical performance at a given age, which is a normative comparison.

Assessment and Diagnosis

How does NIMH's Research Domain Criteria (RDoC) initiative differ from the DSM and ICD?

  • a.It is a dimensional research framework, not a diagnostic manual✓
  • b.It classifies disorders by their response to medication
  • c.It is a categorical manual used for insurance billing
  • d.It replaced DSM categories in U.S. clinical practice

NIMH describes RDoC as a research framework that studies mental health and illness as varying degrees of dysfunction in basic psychological and biological systems, and states that it is not meant to serve as a diagnostic guide or to replace current diagnostic systems. It is therefore neither a billing manual nor a replacement for DSM or ICD, and it is organized by functional domains rather than treatment response.

Assessment and Diagnosis

A client makes many perseverative errors on the Wisconsin Card Sorting Test, continuing to sort by a rule after feedback shows it has changed. Which function is this finding most associated with?

  • a.Set shifting✓
  • b.Receptive vocabulary
  • c.Verbal episodic memory
  • d.Visuoconstructional ability

The Wisconsin Card Sorting Test requires inferring a sorting rule from feedback and shifting when it changes; Milner (1963) linked perseverative errors to frontal-lobe damage, and the test is used as a measure of executive set shifting. It does not primarily measure verbal memory, drawing or construction ability, or word knowledge.

Assessment and Diagnosis

In a county, 200 new cases of depression were diagnosed this year among 10,000 residents who were free of depression at the start of the year. What does this figure describe?

  • a.Incidence✓
  • b.Point prevalence
  • c.Relative risk
  • d.Lifetime prevalence

CDC's Principles of Epidemiology defines incidence as the occurrence of new cases in a population at risk over a specified period, here 200 per 10,000 per year. Prevalence counts all existing cases (new and old) at a point or over a lifetime, and relative risk compares the risk in two groups.

Assessment and Diagnosis

A selection test yields similar mean scores for two groups, but it consistently underpredicts later job performance for members of one group. This is best described as:

  • a.Range restriction
  • b.Differential prediction✓
  • c.Differential item functioning
  • d.Adverse impact

The Standards for Educational and Psychological Testing describe differential prediction as a test-criterion relationship, such as regression slopes or intercepts, that differs across groups, the Cleary (1968) definition of predictive bias. Adverse impact refers to different selection rates, which similar means make unlikely, differential item functioning is an item-level difference for equally able examinees, and range restriction lowers validity coefficients in selected samples.

Report