Peira
Research Brief: A Validated Selection Assessment of Moral and Creative Capacity
For HR, industrial-organizational psychology, and selection professionals. This brief distills the instrument design, validity basis, and defensibility of the platform. The research bibliographies behind the instruments, plus the platform's identity and ZK design notes, are published on this site (linked in §9); the full engineering specification is still internal.
Version 0.1 · 2026-08-11
1. Summary
Selection has a measurement problem, and AI just widened it. Resumes and unstructured interviews no longer separate genuine capability from generated content. Legacy integrity questionnaires are fakeable and carry documented adverse impact. Cognitive aptitude and personality predict part of performance, but not the two dimensions that most often derail or differentiate leaders: moral judgment and creative capacity.
Peira is a research-validated selection assessment built to fill that gap. It pairs four creativity instruments with two morality instruments. All six are structural replicas of peer-reviewed measures. Scoring is deterministic, so every result is reproducible and defensible. The platform aligns with the APA Standards for Educational and Psychological Testing, SIOP guidelines, and the Uniform Guidelines on Employee Selection Procedures (29 CFR §1607), with adverse-impact monitoring built in. It is explicitly non-clinical: selection and descriptive profiling, not diagnosis.
This brief covers what the instruments measure, the validity basis, integrity in the AI era, privacy and compliance, and honest limitations.
2. The problem
Three forces bear on the selection function at once.
AI broke the resume. Candidates generate applications with LLMs at scale. Resumes and unstructured interviews no longer separate genuine capability from generated content.
Legacy integrity tests are fakeable. Likert self-report questionnaires are coached, gamed, and carry documented adverse impact. The U.S. Office of Technology Assessment flagged their core validity threats in 1990: adverse impact, privacy invasion, low fidelity, and faking. The category has barely improved since.
Aptitude is not everything. Cognitive aptitude and personality predict part of performance. They do not predict ethical judgment or creative contribution. Those are the dimensions that derail or differentiate leaders, and most batteries miss them entirely.
3. What it measures
Six instruments across the two domains most selection tools ignore.
| Code | Instrument | Measures | Grounded in |
|---|---|---|---|
| CAI | Creative Achievement | Pro-c / Big-C lifetime creative recognition | Carson, Peterson & Higgins (2005), CAQ |
| AUT | Divergent Thinking | Capacity to generate many, varied, unusual ideas | Guilford (1967); scored via Beketayev & Runco (2016) |
| CPSS | Cognitive Style | Adapter–innovator problem-solving orientation | Kirton (1976) adaption–innovation tradition |
| ECI | Everyday Creativity | Little-c creative practice in daily life | Kaufman & Beghetto (2009), Four-C model |
| MVP | Moral Foundations | Which moral foundations you weigh under conflict | Atari et al. (2023), MFQ-2 |
| MTRS | Dilemma Reasoning | Outcomes vs. duties vs. restraint in hard trade-offs | Gawronski et al. (2017), CNI/ODR model |
Each instrument is a structural replica of a peer-reviewed measure, not an invented scale. Validity is inherited from the source instrument rather than claimed, and each can be pilot-validated against its source's published norms.
4. Validity and professional standards
Every instrument is grounded in decades of published psychometrics. Scoring is deterministic: the same responses always produce the same score, against fixed, published scoring keys. That makes every result reproducible, explainable, and defensible.
Standards alignment.
- APA Standards for Educational and Psychological Testing
- SIOP guidelines on test use and security
- Uniform Guidelines on Employee Selection Procedures (29 CFR §1607)
- Adverse-impact monitoring (4/5ths rule) built in, surfaced to administrators where demographic data exists
- Per-instrument, per-data-category explicit consent (GDPR Art. 9(2)(a))
- Non-clinical framing enforced across consent, results, and reporting
Non-clinical positioning. The platform administers research instruments for organizational selection and descriptive profiling. It is not a medical device, not a diagnostic tool, and not a substitute for professional advice. This framing is enforced in code (a vocabulary lint rule) and in copy, to keep the platform out of clinical-regulatory scope.
5. Integrity: defensible in the AI era
A candidate can copy a test item into a chatbot and paste the answer back. The platform layers defenses against this. No single layer is a silver bullet, so they are stacked.
- Per-candidate perturbation. Every candidate sees a structurally identical, surface-different version of each item. Item-bank leakage and coaching become useless. This protects both validity and audit-defensibility.
- Forced-choice format. No Likert self-endorsement. Candidates judge between two parties with legitimate claims. There is no right answer to fake, and adverse impact runs lower than direct self-report.
- Stylometric and behavioral detection. Free-text responses are screened for the statistical signatures of AI-generated prose. Behavioral signals (timing, paste events, focus) feed a post-hoc classifier. Raw signals never leave the candidate's browser.
- One human, one result. Each result is bound to one unique human via World ID. One test. One person. For life. Every result traces to a verified person, not a bot farm or a duplicate identity.
The output is an integrity band (High / Medium / Low), not a blunt auto-pass/fail. A human reviewer stays in the loop on every flag, and every selection decision stays human-led.
6. Privacy and compliance
Results are keyed to a pseudonymous uniqueness nullifier. Never a wallet. Never a name. Candidates consent explicitly, per instrument and per data category, and grant each organization access on their terms. Data minimization is enforced at the schema, and raw behavioral signals never leave the browser.
Because the platform is non-clinical, response data is positioned as selection-relevant rather than special-category health data, which reduces regulatory exposure for employers and candidates. Every access to candidate data is audit-trailed.
7. Implementation and roadmap
v1: operated platform. A single canonical service: the validated instruments, the scoring engine, and the integrity classifier, administered end-to-end and integrated into your hiring workflow. Canonical instruments are content-committed (their integrity is publicly verifiable), while item plaintext stays gated to preserve validity.
v2: open, verifiable protocol. Canonical instruments become publicly verifiable, any qualified operator can deliver, scoring stays deterministic and reproducible, and candidates gain zero-knowledge selective disclosure of their own results. Governance is upgradeable but timelocked and multisig-controlled, with no silent changes.
8. Limitations (honest)
- Small-cohort reliability. Several instruments (divergent-thinking originality, moral-dilemma parameter estimation) are cohort-relative. At the small cohort sizes realistic for early deployments (5 to 20 candidates), percentiles are noisy. The platform surfaces this caveat, refuses percentile claims below n = 5, and defers cross-organization norming to a later phase.
- Kirton verification gate. Cognitive style (CPSS) is a replica of a commercially licensed measure. Until its citation and licensing are verified directly with the publisher, it ships as a 3-item pilot subset rather than the full 16.
- Adverse impact. Any selection instrument can produce demographic skew. The platform builds in 4/5ths-rule monitoring on opt-in demographic data, treats LLM-detector false positives on non-native English as a real adverse-impact concern, and keeps a human reviewer in the loop on every integrity flag.
- No silver bullet on cheating. The integrity stack raises the cost and detectability of LLM-assisted gaming; it does not eliminate it. The platform says so in candidate and reviewer documentation.
9. References and further reading
Primary sources
- Atari, M., Haidt, J., Graham, J., Koleva, S., Stevens, S. T., & Dehghani, M. (2023). Morality Beyond the WEIRD. JPSP 125(5). DOI 10.1037/pspp0000470.
- Beketayev, K., & Runco, M. A. (2016). Scoring Divergent Thinking Tests by Computer With a Semantics-Based Algorithm. Europe's Journal of Psychology 12(2):210–220. DOI 10.5964/ejop.v12i2.1127.
- Carson, S. H., Peterson, J. B., & Higgins, D. M. (2005). Reliability, Validity, and Factor Structure of the Creative Achievement Questionnaire. Creativity Research Journal 17(1):37–50.
- Gawronski, B., Armstrong, J., Conway, P., Friesdorf, R., & Hütter, M. (2017). Consequences, Norms, and Generalized Inaction in Moral Dilemmas: The CNI Model. JPSP 113(4).
- Kaufman, J. C., & Beghetto, R. A. (2009). Beyond Big and Little: The Four C Model of Creativity. Review of General Psychology 13(1):1–12.
- Kirton, M. J. (1976). Adaptors and Innovators. JAP 61(5). (verification-gated)
Companion documents (published on this site)
- Morality research bibliography: annotated sources behind the moral-reasoning instruments.
- Creativity research bibliography: annotated sources behind the creativity instruments.
- ZK feasibility memo: private, verifiable score credentials for v2.
The implementation plan and instrument-level specification remain internal at this stage.
This research brief describes a research-instrument platform for organizational selection and descriptive profiling. It is not a medical device, does not provide medical or mental health diagnoses, and is not a substitute for professional advice. Selection decisions are made solely by participating organizations. © 2026 Peira. Draft for public review. Feedback welcome.