Methodology & Validity · Evidence Dossier
Every claim, every reference,
every validity note.
This dossier compiles the claim-by-claim verification section of iiop.careercompass/methodology into a single printable document. Each of the nine claims below is presented with its evidence base, peer-reviewed references and the validity caveats that accompany it.
Generated July 29, 2026 · IIOP Career Suite · Institute of Industrial & Organisational Psychology
Contents
- 01 · Authorship & review
- 02 · AERA · APA · NCME
- 03 · IRT scoring
- 04 · Developmental adaptation
- 05 · Reliability α .82–.93
- 06 · Norming & sample
- 07 · Predictive validity
- 08 · Fairness & DIF audit
- 09 · Score-tethered AI Coach
Claim 01 · Authorship & review
Authored and reviewed by chartered I/O psychologists.
Every item is written by a panel of chartered occupational and career-development researchers, then peer-reviewed for face and content validity before it can enter the bank.
Evidence
- 6–9 expert authors per construct with blind peer review.
- Expert-panel Content Validity Index (CVI) ≥ 0.83 per scale.
- Cognitive interviewing with 20+ learners at each developmental band.
References
- Lynn, M. R. (1986). Determination and quantification of content validity. Nursing Research.
- Willis, G. B. (2005). Cognitive Interviewing: A Tool for Improving Questionnaire Design.
Validity note
Content validity is a necessary but not sufficient condition; it is triangulated with internal-structure and criterion evidence in the sections below.
Claim 02 · AERA · APA · NCME
Aligned with AERA · APA · NCME (2014) and ITC guidelines.
Development, adaptation and reporting follow the internationally recognised Standards for Educational and Psychological Testing and the ITC guidelines for adapted testing and computer-based assessment.
Evidence
- Development lifecycle mapped 1:1 to the 12 Standards clusters.
- Translations follow ITC forward-translation, reconciliation and back-translation.
- Reporting conforms to EFPA Test Review Model criteria and ISO 10667.
References
- AERA, APA & NCME (2014). Standards for Educational and Psychological Testing.
- International Test Commission (2017). ITC Guidelines for Translating and Adapting Tests (2nd ed.).
- EFPA (2013). Test Review Model, v4.2.6.
- ISO 10667-1/2:2020 — Assessment service delivery.
Validity note
Alignment is a process claim, not a certification. Independent auditors are welcome under NDA to review our compliance matrix.
Claim 03 · IRT scoring
IRT-based scoring, not raw sum-scores.
Responses are scored via 2-parameter logistic and Graded-Response IRT models, converted to age- and country-normed T-scores (M = 50, SD = 10) and percentiles.
Evidence
- All calibrated items report discrimination (a) and difficulty (b) parameters.
- Items with a < 0.70 or large drift on annual re-analysis are retired.
- Composite indices (CDSE, PsyCap, Protean) combine scales with published regression weights.
References
- Samejima, F. (1969). Estimation of latent ability using a response pattern of graded scores. Psychometrika Monograph.
- Embretson, S. E. & Reise, S. P. (2000). Item Response Theory for Psychologists.
- de Ayala, R. J. (2009). The Theory and Practice of Item Response Theory.
Validity note
IRT assumes unidimensionality within a scale; multidimensional composites are estimated only where CFA supports the structure (CFI ≥ 0.94, RMSEA ≤ 0.06).
Claim 04 · Developmental adaptation
Developmentally adaptive — one spine, four expressions.
Reading level, item format, session length, coach persona and interpretation vocabulary evolve across Primary, Secondary, Higher Ed and Professional tiers while the underlying constructs remain psychometrically equivalent.
Evidence
- Configural, metric and scalar measurement invariance tested across tiers and countries.
- Within-tier graded-response CAT halts once SE(θ) ≤ 0.30.
- Reading level: Grade 3 pictorial → Adult professional.
References
- Vandenberg, R. J. & Lance, C. E. (2000). A review and synthesis of the measurement invariance literature. Organizational Research Methods.
- van der Linden, W. J. & Glas, C. A. W. (2010). Elements of Adaptive Testing.
Validity note
Full scalar invariance is not always achieved between the youngest (Primary) and adult tiers; partial invariance is documented in the Technical Manual and cross-tier comparisons are flagged accordingly.
Claim 05 · Reliability α .82–.93
Reliability α .82–.93 across scales.
Every published scale reports Cronbach’s α, McDonald’s ω and 4-week test–retest correlations in the Technical Manual.
Evidence
- Cronbach’s α: .82–.93 across published scales.
- McDonald’s ω reported alongside α for congeneric reliability.
- Test–retest r = .78–.86 at a 4-week interval.
References
- Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika.
- McDonald, R. P. (1999). Test Theory: A Unified Treatment.
- Revelle, W. & Condon, D. M. (2019). Reliability from α to ω. Psychological Assessment.
Validity note
Alpha is a lower bound and is sensitive to scale length; ω is the primary reliability index for short adaptive scales. Values are norm-sample specific — not fixed constants.
Claim 06 · Norming & sample
27,800+ respondents across 18 countries.
Norms are stratified by age band, gender, country and school type. Country-specific tables are used whenever the in-country N ≥ 400.
Evidence
- Career Explorer Junior™ N = 4,800 · Career Explorer™ N = 9,600.
- Career Navigator™ N = 8,200 · Career Reinvention™ N = 5,200.
- Re-norming cadence: 18–24 months per instrument.
References
- AERA, APA & NCME (2014). Standards, Ch. 5 (Reference populations).
- IIOP (2025). Career Suite Norming Census, v2.1 — available on request.
Validity note
Where in-country N < 400, the closest cultural-cluster norm is applied and every affected report carries an explicit norming-cluster flag.
Claim 07 · Predictive validity
Predicts real career outcomes — not just self-report.
The Suite predicts subject-choice satisfaction, university major-fit and 6-month career-clarity gains in a 3,100-student longitudinal cohort.
Evidence
- Subject-choice satisfaction: β = .41.
- University major-fit: β = .38.
- 6-month career-clarity gain: Cohen’s d = 0.72.
- Convergent r ≥ .55 with RIASEC / HEXACO-60 / CAAS-International.
References
- Savickas, M. L. & Porfeli, E. J. (2012). Career Adapt-Abilities Scale. Journal of Vocational Behavior.
- Rounds, J. & Su, R. (2014). The nature and power of interests. Current Directions in Psychological Science.
Validity note
Effect sizes reflect the current longitudinal cohort and will vary across populations; predictive coefficients are re-estimated each re-norming cycle.
Claim 08 · Fairness & DIF audit
Audited for bias and fairness every year.
Differential Item Functioning is tested annually by gender, age band, country and first-language using Mantel–Haenszel and IRT-LR procedures.
Evidence
- Items flagged large-DIF are retired or re-parameterised before release.
- Accessibility follows WCAG 2.2 AA, including screen-reader and keyboard flows.
- Coach outputs are audited on a rotating sample for stigmatising language.
References
- Holland, P. W. & Wainer, H. (1993). Differential Item Functioning.
- Zumbo, B. D. (2007). Three generations of DIF analyses. Language Assessment Quarterly.
- W3C (2023). Web Content Accessibility Guidelines 2.2.
Validity note
Fairness testing detects statistical DIF, not lived experience of bias; qualitative review with representative learners complements the quantitative audit.
Claim 09 · Score-tethered AI Coach
The AI Coach is score-tethered, not score-invented.
Career, subject and university matches are computed by the deterministic scoring engine before any language model is called. The coach explains matches — it does not generate them.
Evidence
- System prompts constrain the coach to interpretive thresholds present in the profile.
- Deterministic recommendation core runs before LLM narration.
- Low-confidence and safeguarding-flagged sessions route to a chartered psychologist.
References
- APA (2023). Guidelines for the use of AI in psychological practice.
- NIST AI Risk Management Framework 1.0 (2023).
Validity note
AI narration is a communication layer, not a measurement layer. Recommendations remain reproducible from the scored profile without the coach.
For the full psychometric derivation, calibration tables and re-norming schedule, request the IIOP Technical Manual from the Methodology page.