Methodology

Last updated: 11 September 2026

How the assessment is built, how it is scored, and what has not been validated yet. The framework it measures against has four decades of research behind it; the question set is ours and does not. This page keeps those two apart, because conflating them is the most common overclaim in this category.

The framework is not ours. The questions are.

The assessment measures against the Competing Values Framework, which came out of research by Robert Quinn and John Rohrbaugh in the early 1980s and was developed into a culture instrument by Kim Cameron and Robert Quinn. It is one of the most widely applied culture models in organisational research, with four decades of literature behind it and thousands of published studies using it. It describes culture on two axes — internal versus external focus, and flexibility versus control — which produce four orientations: Clan, Adhocracy, Market and Hierarchy.

  • Quinn, R. E., & Rohrbaugh, J. (1981). A competing values approach to organizational effectiveness. Public Productivity Review, 5(2), 122–140.
  • Quinn, R. E., & Rohrbaugh, J. (1983). A spatial model of effectiveness criteria: towards a competing values approach to organizational analysis. Management Science, 29(3), 363–377.
  • Cameron, K. S., & Quinn, R. E. (2011). Diagnosing and Changing Organizational Culture: Based on the Competing Values Framework (3rd ed.). Jossey-Bass.
  • The distinction that matters: the framework is established and independently validated. The specific question set in this product was written by us and has not been. Those are two different things, and this page keeps them apart throughout.

What the candidate actually does

Twenty blocks. Each describes an ordinary situation at work and offers four responses, one written against each of the four orientations. The candidate puts all four in order of preference. There are no free-text answers, no open questions, and no neutral option — the ranking is complete or the block is not answered. The order the four responses appear in is shuffled, so a candidate who always picks the first option produces no usable profile rather than a Clan-heavy one.

  • Forced ranking rather than a rating scale, deliberately. Asked to rate, people rate almost everything as important; the result is a flat profile that distinguishes nobody. Ranking forces the trade-off that is the actual signal.
  • It also blunts two well-known response biases: acquiescence, the tendency to agree with whatever is presented, and social desirability, answering as the person you would like to be seen as. When every option is a reasonable thing to prefer, there is no obviously correct answer to give.
  • The cost of that choice is real and is stated here rather than buried: ranked scores are relative within one person. They say which orientation someone leans toward compared with their own other three, not how strongly they lean compared with another candidate. Comparisons between two people are therefore weaker than a rating scale would give, and that is the trade this design accepts.
  • Response time per block is recorded, which is what makes it possible to identify a completion too fast to have been read.

How the questions were written, and what constrains them

Every item was written by hand for this product. None is drawn from a published question bank, and that is enforced rather than asserted: a test hashes all 2,400 strings across the twelve languages against a checked-in list of digests from external instruments, so a copied question fails the build. The plaintext of that list is deliberately absent from the repository, so the guard cannot itself become the leak.

  • Within each block, the four options are held to a similar length. An option noticeably longer than its siblings reads as the considered, "correct" answer and attracts rankings on presentation alone.
  • Across the twenty blocks, option length is held to a similar spread per orientation, so no single orientation is systematically the wordiest — which would bias the whole instrument, not just one block.
  • The company-facing and candidate-facing banks are deliberately not mirror images of each other. An employer describing its own culture and a candidate describing their preferences are answering different questions, and identical wording would invite the employer to answer as the culture it aspires to.
  • All of the above are checked by tests that run in the ordinary suite, so a content regression fails on a fresh clone with nothing configured.

What it deliberately does not measure

The instrument was audited item by item, before launch, against every special category of personal data in Article 9 of the GDPR and against Article 10. It collects and infers none of them. That result is held in place by 74 tests which assert the absence rather than describing it, so an item added later that asks about any of these fails the build.

  • No question asks about health, disability, pregnancy, sex life or sexual orientation, racial or ethnic origin, religious or philosophical belief, political opinion, trade union membership, genetic or biometric data, or criminal record.
  • It measures cultural preference. It does not measure competence, ability, intelligence, personality, integrity, resilience or suitability for a role, and no output of this product should be read as measuring any of them.
  • It is not a predictor of job performance. No claim is made that a high congruence score leads to a better hire, because establishing that would require a longitudinal study this product has not run.
  • Nothing in the product accepts or rejects anyone. There is no pass mark, no threshold, no shortlist and no reject action, and the terms of service require that a result is never the sole basis for an employment decision.

The scoring is arithmetic, and it is published

A ranking is converted to points by a fixed formula: within each block the first-ranked option scores 3, then 2, then 1, then 0, and the points accumulate to the orientation that option was written against. The four totals are normalised to sum to exactly 100, which is what makes a profile readable as percentages. The same rankings always produce the same numbers — nothing is sampled, nothing is estimated, and no model is involved in producing any figure.

  • Fit between a candidate and a company is computed congruence between two profiles that each sum to 100. It is a similarity measure between two descriptions, not a probability of anything.
  • Because it is deterministic, a score quoted to a client today is the score they see next quarter. This is a genuine difference from any scoring that runs through a language model, where the same input can produce a different number.
  • The dimension readings — how someone leans under pressure, on trust, on ambition and on risk — are computed the same way from the same rankings. They are descriptive summaries of the answers, not separate tests.

Where the language model is, and where it is not

A language model writes the readable prose in the report. It does that after every number has already been calculated, and it receives only those numbers. It is never given a name, an email address, an individual answer, or any note a recruiter has written. It computes nothing, and if it were removed the scores would be unchanged.

  • This is stated on the privacy notice and in the terms as well, because it is the single thing people most reasonably worry about when a product mentions AI.
  • The prompts carry an explicit vocabulary restriction: the model is forbidden from describing a person in terms of resilience, stress tolerance, coping, wellbeing, burnout, communication style, cultural or national characteristics, attitude to authority, or integrity — none of which this instrument measures, and several of which would stray into special-category inference.
  • The written summary is decision support. The deterministic scores are the authoritative output.

What has not been established

This is the section a buyer should read first, and it is here in full rather than summarised. The framework this instrument measures against is validated. This instrument is not. Everything below is a real gap, stated because a reader who finds it themselves after reading a confident claim has learned something worse about us than the gap itself.

  • There is no published validity study for this question set. No convergent validity against an established measure, no discriminant validity, no criterion validity against any workplace outcome.
  • There is no reliability estimate. No test-retest coefficient has been computed, because that requires the same respondents taking it twice with a suitable interval.
  • There is no norm group. Scores are not compared against a reference population, and no percentile is reported, because there is no population to compare against.
  • No Cronbach’s alpha is quoted, and that is a methodological point rather than an omission. Because the four scores are forced to sum to 100 they are linearly dependent, which drives the average correlation between them negative — approximately −1/3 for four scales — no matter how well the items are written. Classical internal-consistency and factor-analytic statistics are not valid on this kind of data. A forced-choice instrument advertising an alpha is either quoting it from a separate rating-scale form or quoting a number that does not mean what it appears to mean.
  • The twelve languages are not equally verified. The English set is the source; the others are machine translations, and the product does not claim a language as available in public copy until a fluent human has reviewed it.

How this gets validated, and what each stage needs

Published in advance so that it is a commitment rather than a description written after the fact. The item-level data required for all of it is already retained on every completed assessment — the rank of each option in each block, the orientation it belongs to, the exact text the respondent read, and the time they spent — so the analyses below become possible as the number of completed assessments grows, without any change to what is collected.

  • Stage one, from roughly thirty completed assessments and needing no statistician: completion and abandonment rate by block, time on task, and the distribution of scores across the four orientations. This catches the practical failures that matter most and are cheapest to fix — a block everyone gives up on, wording nobody understands, or an orientation that never places first because its options are written less appealingly than the rest.
  • Stage two, from roughly two hundred to three hundred completed assessments: item-level analysis using a model appropriate to forced-choice data. The established approach is Thurstonian item response theory (Brown & Maydeu-Olivares, 2011), which estimates item parameters and tests whether the four-orientation structure holds, without the invalidity that classical factor analysis has on ranked data. This needs a psychometrician, not a developer.
  • Stage three, and the most persuasive study available: convergent validity. Administer a published instrument measuring the same framework alongside this one to a subsample and report the correlation. If the two agree, this question set measures what it claims to; if they do not, that is worth knowing before a customer discovers it.
  • Stage four, criterion validity — whether cultural fit predicts retention or satisfaction — requires following placements over years with agency cooperation. It is listed for completeness and is not planned.
  • The intended route for stages two and three is a research partnership with a university psychology faculty rather than paid consultancy: a supervisor needs publications and a student needs a dataset, and a novel forced-choice instrument on a well-known framework is a publishable subject. Anonymised item-level responses carry no identifying data, so such a partnership needs no candidate consent and no transfer of personal data.
  • This page will be updated with results as each stage completes, including results that are unfavourable.