This evidence pack reports psychometric analyses of the full AQme global dataset, conducted as part of AQai's 2026 research cycle. It is the technical annex behind the AQme Technical & Scientific Manual, written for reviewers who want the numbers, the methods, and the limitations stated plainly.
Analyses were performed on the anonymised scored dataset: subscale, dimension and index scores with demographics, organisation identifiers and assessment dates. Scores are analysed at subscale level (0–100 display scale). Version 2.0 adds the item-level programme: the full platform response export (20,187 completions, 93 scored items) enables refreshed reliability with McDonald's omega, item-level confirmatory factor analysis, and item-level differential functioning screens, reported in Section 9. Every analysis here is scripted and reproducible, and will be re-run on the growing dataset at each recalibration cycle.
| Analysis | Method | Section |
|---|---|---|
| Norm base composition & temporal stability | Descriptives by year, region, gender, age, industry, job level | 2 |
| Reliability & measurement precision | Internal consistency (2022, N=4,849); reassessment stability and SEM (n=606 repeat pairs) | 3 |
| Dimensionality & structure | EFA (promax), competing CFA models (ML), discriminant single-factor test | 4 |
| Structural invariance across groups | Per-group CFA, loading congruence (Tucker's φ) across gender, region, age | 5 |
| Fairness screen | Group effect sizes (Cohen's d); differential test functioning screen (group beta controlling total score) | 6 |
| Nomological network | Theory-predicted inter-scale correlations vs observed | 7 |
| Norm-referenced banding | Verification of tercile method; threshold history 2021–2025 | 8 |
16,491 unique working adults completed a first AQme assessment between January 2020 and December 2025. Annual volumes grew from 1,096 (2020) to 4,753 (2025).
| Characteristic | Composition |
|---|---|
| Geography | 141 countries. North America 48.2%, Europe 27.8%, Asia-Pacific 13.9%, Africa/Middle East 4.9%, Australasia 3.3%, Latin America 1.9% |
| Gender | Male 48.8%, female 46.4%, non-binary/transgender/undisclosed 4.8% |
| Age | 18–24: 4.7% · 25–34: 21.3% · 35–44: 31.6% · 45–54: 28.2% · 55+: 14.2% |
| Work profile | 18 industry categories; job levels from entry to executive (experienced specialists 24%, middle management 22%, first-level management 19%, senior/executive 15%, entry-level 9%, other 11%) |
Temporal stability. Annual population means for overall AQ range 601–617 (grand mean 608, SD 93) — a band of ±1.5% across six collection years, with subscale means similarly stable. The benchmark population is not drifting materially, which supports both the comparability of reports across years and the rolling-window recalibration design (Section 8).
Distributional properties. Most subscales are approximately normal with mild negative skew. Two show marked skew and ceiling concentration — Grit (skew -1.28; 4.3% of all subscale scores sit at the display ceiling, concentrated in Grit and Team Support) — a known property of self-report perseverance measures, handled in banding by subscale-specific thresholds (Section 8) and flagged in coach guidance.
Internal consistency was estimated on 4,849 completions (2022 cycle; refresh with McDonald's omega scheduled with item-level data). Reassessment stability was computed from the 606 users with two or more assessments, correlating prior with current scores over naturally varying intervals; SEM was derived from the distribution of retest differences (SDdiff/√2), which converges with the alpha-based estimate where both are available.
| Scale | α (2022) | Stability r | SEM (pts) | 68% band (±) | 95% band (±) |
|---|---|---|---|---|---|
| Grit | .71 | .80 | 6.5 | 6.5 | 12.7 |
| Mental Flexibility | .87 | .78 | 8.7 | 8.7 | 17.1 |
| Mindset | .79 | .79 | 8.3 | 8.3 | 16.3 |
| Resilience | .79 | .81 | 9.1 | 9.1 | 17.8 |
| Unlearn | .67 (beta n=211) | .79 | 8.5 | 8.5 | 16.7 |
| Emotional Range | .68 (2-item) | .81 | 10.9 | 10.9 | 21.4 |
| Extraversion | .70 (2-item) | .84 | 11.8 | 11.8 | 23.1 |
| Hope | .85 | .81 | 7.9 | 7.9 | 15.5 |
| Motivation Style | .75/.76 (poles) | .74 | 5.8 | 5.8 | 11.4 |
| Thinking Style | .72 | .76 | 9.6 | 9.6 | 18.8 |
| Company Support | .86 | .81 | 11.2 | 11.2 | 22.0 |
| Emotional Health | .78 | .78 | 9.6 | 9.6 | 18.8 |
| Team Support | .80 | .82 | 8.7 | 8.7 | 17.1 |
| Work Environment | .83 | .79 | 8.9 | 8.9 | 17.4 |
| Work Stress | .90 | .85 | 12.7 | 12.7 | 24.9 |
Reading. Every subscale shows reassessment stability of .74 or above — strong for measures that deliberately include state-sensitive content. The three scales with marginal alpha have stability of .79–.84, indicating the alpha values (two-item facets; beta-phase subsample) understate their actual consistency. The SEM columns are the practically important numbers: a subscale score should be interpreted as a range (roughly ±6–13 points at 68% confidence), band boundaries should be treated as zones rather than lines, and reassessment change is meaningful when it exceeds the SEM.
Question: do the 15 displayed subscales measure distinct things, and how do they organise? Method: exploratory factor analysis (promax) and maximum-likelihood confirmatory models on standardised subscale scores, N = 15,112 first assessments with complete data. Note the level of analysis: these are subscale-level models. They test how the 15 measured constructs relate — not item-level measurement, which is the next research step.
A single-factor model fails badly (CFI = .68, RMSEA = .13), and mean inter-scale correlations are moderate (within-dimension |r| = .34, between = .24, maximum r = .72). Adaptability as measured by AQme is genuinely multidimensional: the instrument is not re-measuring one construct fifteen times, and no subscale is redundant with another.
EFA supports three factors (eigenvalues 5.05, 1.71, 1.19): an emotional-wellbeing cluster (Emotional Health, Emotional Range, Resilience, Mindset), an organisational-environment cluster (Work Environment, Team Support, Company Support), and a cognitive-flexibility cluster (Mental Flexibility, Unlearn, Thinking Style, Motivation Style, Hope). The three-factor structure confirms that person-level capability, emotional functioning, and situational context are distinguishable data clusters — the core premise of a multi-domain model.
A reflective CFA forcing subscales onto their A.C.E. assignments fits below conventional thresholds (CFI = .80, RMSEA = .11), as does the best empirical alternative at this level of aggregation (CFI = .84, RMSEA = .10). Two reasons, both by design. First, Character was never conceived as a unitary trait: its subscales are deliberately heterogeneous preference continua (its loadings — Motivation Style .39, Thinking Style .58 — reflect that design, and a reflective factor is the wrong model for it). Second, theory predicts cross-domain links (resilience with emotional range; stress with wellbeing) that a simple-structure model forbids. Accordingly, AQai positions A.C.E. as an organising framework for interpretation — how, why, when — and reports dimension scores as documented weighted composites (formative indices), not as estimates of three latent traits. Item-level CFA, including bifactor and ESEM alternatives, is scheduled with the item-level dataset.
Does the instrument behave the same way for different populations? The A.C.E. measurement model was fitted separately within gender, region and age subgroups; factor-loading profiles were compared with the pooled solution using Tucker's congruence coefficient (φ ≥ .95 conventionally indicates factorial similarity).
| Group | n | Loading congruence φ |
|---|---|---|
| Male | 7,785 | .999 |
| Female | 7,327 | .999 |
| North America | 7,729 | 1.000 |
| Europe | 4,249 | .999 |
| Asia-Pacific | 2,260 | .989 |
| Under 35 | 4,047 | .999 |
| 45 and over | 6,773 | .999 |
Loading profiles are effectively identical across all groups tested (φ = .989–1.000), with per-group model fit closely matching the pooled fit. The relationships among the subscales — what goes with what — are the same for men and women, across major regions, and across age bands. Formal multi-group invariance testing (configural/metric/scalar) at item level is scheduled with the item dataset.
Scale-level group differences are small throughout the norm base. The largest gender effect anywhere in the instrument is Emotional Range (d = 0.32, women reporting higher emotional responsiveness — the direction found across the Big Five literature); overall AQ differs by d = 0.11. Regional means span roughly a quarter of a standard deviation. Age shows the theoretically expected positive gradient (overall AQ 589 at 18–24 rising to 627 at 55+), consistent with published findings on adaptability and experience.
A differential-functioning screen — regressing each subscale on group membership while controlling for overall AQ — shows most subscales within ±0.15 SD at equal total score. Three exceed it modestly (Mindset +0.19 for women; Emotional Range -0.25; Emotional Health -0.17), a pattern consistent with substantive gender differences in affect reporting rather than measurement artefact — but distinguishing those explanations rigorously requires item-level DIF analysis, which is scheduled. No group difference approaches the magnitude that would concern a reviewer under EEOC-style adverse-impact framing, and AQme is in any case not used for selection decisions.
If the subscales measure what they claim, their inter-correlations should reproduce known relationships from the literature. They do:
| Predicted relationship | Basis | Observed r |
|---|---|---|
| Resilience ↔ Emotional Range (positive) | Resilience–emotional stability link (BRS literature) | +.55 |
| Emotional Health ↔ Work Stress (negative) | Overwork erodes thriving (JD-R literature) | -.38 |
| Team Support ↔ Work Environment (positive, strong) | Psychological safety underpins learning climate (Edmondson) | +.72 |
| Team Support ↔ Company Support (positive, distinct) | Related but separable support constructs (POS literature) | +.50 |
| Mental Flexibility ↔ Unlearn (positive) | Cognitive flexibility cluster | +.53 |
| Mindset ↔ Hope (positive, distinct) | Optimism and hope related but separable (Snyder; Scheier & Carver) | +.45 |
| Grit ↔ Hope (positive) | Agency–perseverance link | +.47 |
| Extraversion ↔ Emotional Range (low) | Independent Big Five facets | +.29 |
Every prediction is met in direction and approximate magnitude, and the two strongest correlations (.72, .55) sit exactly where theory puts them. Combined with the age gradient (Section 6), this is consistent convergent and discriminant evidence that the subscales measure their intended constructs.
Note on composite indices: the Change Readiness Index (CRi) and Reskill Index (RSi) are documented, weighted collections of the measured subdimensions, maintained under the versioned scoring specification; their measurement quality inherits from the component scales reported above.
Band thresholds (low/medium/high) are defined as approximate terciles of the global distribution per subscale, recalibrated against a rolling one-year window — a system introduced in April 2022 with versioned reports and participant/partner communication. Verification against the current dataset confirms live thresholds track empirical terciles within a few points on most subscales (e.g. Grit cuts ≈74/85 vs empirical terciles 76/86). The complete threshold history across six recalibration periods (March 2021 – March 2026) shows movement of at most a few points per scale (e.g. Grit low cut ranging 73–76, high cut 85–89; most cuts within ±3), demonstrating both that recalibration matters and that the benchmark is stable and continuously maintained.
Two methodological refinements are scheduled: moving from the normal-approximation cut calculation to direct empirical percentiles (removing a distributional assumption that mis-places cuts slightly on skewed scales such as Grit), and publishing the live threshold table to administrators at each recalibration.
The full platform response export (20,187 completions; 93 scored items; EU and US instances) enables the definitive item-level programme. Response completeness is very high: 99.3% of completions answered 80+ scored items, reflecting the conversational format's enforcement of responses — partial records are rare.
| Scale | n | α | ω / SB | Reading |
|---|---|---|---|---|
| Work Stress | 20,181 | .89 | .89 | Excellent |
| Company Support | 19,985 | .84 | .85 | Good |
| Mental Flexibility | 20,039 | .84 | .84 | Good |
| Change Uncertainty | 20,054 | .83 | .84 | Good |
| Hope | 20,178 | .83 | .83 | Good |
| Work Environment | 20,046 | .81 | .81 | Good |
| Explore & Transform | 19,994 | .80 | .81 | Good |
| Supplementary research items (task-adaptability self-ratings; not a reported dimension) | 20,045 | .79 | .80 | Research use only |
| Resilience | 20,171 | .77 | .78 | Acceptable |
| Emotional Health (re-keyed) | 20,049 | .77 | .78 | Acceptable — see 9.2 |
| Team Support | 19,980 | .77 | .77 | Acceptable |
| Utilize & Improve | 20,047 | .76 | .77 | Acceptable |
| Mindset | 20,041 | .75 | .76 | Acceptable |
| Motivation Style — Play to Protect pole | 20,050 | .74 | .76 | Acceptable |
| Motivation Style — Play to Win pole | 20,050 | .74 | .75 | Acceptable |
| Grit | 20,042 | .72 | .74 | Acceptable |
| Thinking Style | 20,052 | .70 | .70 | Acceptable |
| Extraversion (2-item) | 20,052 | .69 | .70 (SB) | Marginal — revision underway |
| Emotional Range (2-item) | 20,051 | .66 | .67 (SB) | Marginal — revision underway |
| Unlearn | 15,540 | .64 | .64 | Marginal — revision underway |
The full-sample estimates track the 2022 analysis closely (most within ±.03), confirming the reliability profile is a stable property of the instrument rather than a sample artefact. The Unlearn estimate is now definitive (n = 15,540, not the earlier beta subsample): at .64 the scale is confirmed as the bank's weakest, and a structured item-improvement programme for Unlearn and the two two-item facets is underway (all item-rest correlations on Unlearn fall between .31 and .46 — the scale needs new items, not pruning).
Item-level analysis located the source of the Emotional Health scoring anomaly: the two negative-affect items are not reverse-keyed at capture, and the historical weighted-average compensated for this. Re-keying the two items yields a conventional scale (α = .77, all inter-item correlations correctly signed) and permits standard mean scoring. The correction is specified for the next versioned scoring release.
Confirmatory models at item level were estimated with both maximum likelihood and the ordinal-appropriate DWLS estimator (n = 6,000 subsample). Under DWLS — the correct treatment for 7-point ordinal responses — the correlated-factor models within each master dimension meet conventional fit thresholds: Ability five-factor CFI = .94, RMSEA = .055; Character six-factor CFI = .94, RMSEA = .054; Environment five-factor CFI = .98, RMSEA = .045. Per-subscale unidimensional models fit at CFI ≥ .95 on all multi-item scales (Mindset .948, reflecting the documented method variance of positively/negatively keyed items). The item-level measurement model is confirmed: each subscale measures one thing, and the subscales are distinct but related within their dimensions. ML results (CFI .81–.91, RMSEA .063–.069) are retained for comparability and behave as expected for ordinal data under ML.
Screening all 93 items for differential functioning between the EU and US platform populations (item regressed on group controlling rest-score): 86 of 93 items fall within ±0.15 SD. Seven items show modest differential functioning (|0.15–0.23|), concentrated in Mental Flexibility (three of nine items) and single items in Grit, Mindset, Unlearn and Play to Win — flagged for item-revision review. None reaches a magnitude that materially shifts scale scores. Demographic (gender, age) item-DIF requires linking response records to profile data and is scheduled with the platform team.
Completing the scoring methodology review, subscale scores were recomputed from raw items under the current truncation rule (floor) versus conventional rounding (round-half-up). The choice changes the assigned band for 6.8% of subscale scores overall, and for 17.6% on Grit, where the score distribution is densest near the band cuts — truncation is not cosmetic. Quantified impact tables support the specified change to round-half-up (applied once, at display) in the next versioned scoring release, alongside the Emotional Health re-keying (9.2) and documentation of the display bounds.
Data-quality indicators relevant to the conversational administration format: straightlining (within-person SD < 0.5 across ~90 items) occurs in 0.27% of completions; extreme-maximum response styles (≥80% of items at the top option) in 0.27%; the median within-person response SD is 1.46 scale points, indicating healthy differentiation across items. Response completeness is 99.3% at 80+ items. These are strong engagement indicators for a ~90-item self-report instrument; a controlled comparison against non-conversational administration remains the definitive method-effects test and stays on the research roadmap.
Reproducibility. All analyses were scripted (Python: pandas, factor-analyzer, semopy) against the anonymised export of 1 January 2026 and are re-runnable at each recalibration cycle. Method queries: hello@aqai.io