Resquites High-Range Tests

THETIS · engine 2.0 · bank 1.1

THETIS technical manual

How THETIS is built and scored, and how precisely it measures: the item bank, the item response model, the adaptive algorithm, the stopping rule, the score report, and a simulation study of 10,000 virtual candidates on the live bank. Updated 2026-10-06.

1. Purpose and intended use

THETIS estimates an adult's fluid reasoning ability on the deviation-IQ scale (mean 100, SD 15), for self-knowledge, research and admission to the Resquites high-IQ societies. It is an online reasoning test, not a clinical or diagnostic assessment, and it should not be used for clinical, legal, educational-placement or employment decisions. It is intended for ages 16 and up; norms are being built for ages 16 to 75.

2. Construct

THETIS measures abductive–dynamic reasoning: inferring a hidden rule from evidence (abduction), testing it against every observation, and updating it when the evidence shows the rule has changed (dynamic updating). Every item has one keyed answer, the prediction of the simplest explanation of everything shown, verified by computer search of the item's full hypothesis space; every distractor is the prediction of a modeled reasoning error. See About THETIS.

3. Item bank

Bank 1.1 holds 12,597 scored items in seven families (the 7 public practice items are excluded from scoring). Items have 4 options (1,670 items), 6 options (10,767 items), 8 options (160 items). The table shows the items by family and by the IQ at which each is most informative (its difficulty b on the IQ scale).

IQ bandMSCTLGOTotal
40–70110901890016110515
70–85131200452401624100756
85–10019676009600368
100–115118237402150150242491,348
115–1302582133363001501512581,666
130–1455106324624503005163113,181
145–16024724802602603884431,846
160–17500004204013231,144
175–1900003203033224241,369
190–240+0000337661404

Families: M = Mutable Matrices; S = Drifting Strings; C = The Oracle's Ledger; T = Translations; L = Levers and Lamps; G = Turnstiles; O = Orbits.

The bank is densest from IQ 115 to 190 and thinnest from IQ 85 to 100 (368 items, from three families). The routing block draws its first five items from IQ 85 to 115 (1,716 items), which is enough for exposure control. The generators behind the bank have produced over 100,000 machine-verified candidate items for five families; the next bank release adds items in the IQ 70 to 115 range first, where the specification sets the largest targets.

4. Item model

The probability of a right answer is the four-parameter logistic model: P(θ) = c + (d − c) / (1 + e−a(θ − b)), with θ the ability on the logit scale (IQ = 100 + 15θ under the current norms). For every item a = 1.8 and c = 1 / (number of options) (0.125, 0.167, 0.250); b comes from the PROTEUS complexity model. The upper asymptote d = 1 − slip allows for lapses: slip = .10 on the first five items, while the format is new, and .02 afterwards (Barton & Lord, 1981; Rulison & Loken, 2009). These parameters are a priori and provisional until recalibrated from human answers.

5. Adaptive algorithm

  • Routing block. Items 1 to 5 come from the IQ 85–115 band (b between −1 and +1).
  • Mastery gate. No item above IQ 130 until three right answers to items at IQ 85 or harder.
  • Crash protection. The routing estimate may fall at most 0.5 logit per item; the reported score always uses the full estimate.
  • Selection. Maximum Fisher information at the routing estimate, chosen at random among the 16 (first three items) or 8 most informative eligible items (randomesque exposure control).
  • Content balance. Families compete on information; among those within 40% of the best, the family furthest below its equal share is chosen (Kingsbury & Zara, 1989), never the same family twice in a row when another is eligible.
  • No repeats. A person never sees an item twice, across all their sessions, nor two items built on the same structure in one session.

6. Stopping rule and scoring

Ability is the expected a-posteriori (EAP) estimate on a grid from θ = −6 to 14 (step .05) with a normal prior N(0, 3²), wide enough that very high scores are reachable. The test stops when the posterior standard deviation reaches 0.30 (4.5 IQ points) after at least 20 items, or after 45 items. The report gives the IQ with its 95% range (±1.96 SE), the percentile and, above IQ 130, the rarity. THETIS is untimed: response times never enter the score.

7. Score report

Besides the IQ and its range, the report gives seven facet scores (augmented subscores from the families' items), a confidence-calibration index from the optional confidence rating given with each answer, estimated equivalents on other scales (for orientation only), a note on what THETIS measures, and the provisional-norms disclaimer; scores above IQ 145 and below IQ 70 carry the extreme-range and floor disclaimers. Each finished session has a PDF certificate with an ID and QR code, verifiable at /verify/.

8. Simulation study

10,000 virtual candidates in the specification's seven tiers, each with a personal attention-slip rate drawn between .02 and .10, took THETIS through the server's own engine on the live bank (seed 20261006). Their answers follow the item model with their own slip rate, which is larger than the engine assumes after item 5: the results include that misfit.

TierTrue IQNItems (min–max)BiasRMSEMean SE95% coverageStopped at 45
Clinical40–7050039.0 (28–45)−0.24.754.9194.8%31.0%
Low average70–851,00033.2 (26–45)−1.14.744.4694.3%1.7%
Average85–1154,00031.3 (24–45)−1.04.864.4592.8%1.7%
High average115–1302,00030.8 (24–45)−1.25.084.4491.8%1.1%
Gifted130–1601,50031.4 (26–45)−1.85.064.4491.6%0.9%
Profoundly gifted160–20080035.0 (28–45)−2.66.064.5088.8%7.5%
Beyond200–24020041.8 (33–45)−5.612.935.2275.0%47.0%

Bias and RMSE in IQ points (estimate minus true). Coverage: share of 95% ranges containing the true IQ. The routing rule held in 100.0% of sessions (all first five items within IQ 85–115).

Across all tiers the correlation between true and estimated IQ is .99 and the root-mean-square error 5.30 IQ points, with 32.3 items on average. Estimates run slightly low (about 1 to 2 points in the middle and upper tiers) because the virtual candidates slip more often than the engine's .02 assumes; human data will show which slip rate fits real candidates. Coverage of the 95% range is a little under 95% for the same reason.

9. Reliability and precision

  • Marginal reliability in a normal population sample of 2,000 virtual candidates: .92 (correlation of true and estimated IQ .95), with 31.6 items on average.
  • Retest with entirely new items (1,000 candidates, two sessions each): r = .90; mean change −0.20 IQ points.
  • Standard error at the end of a session: 4.5 IQ points or less unless the 45-item limit is reached, so the 95% range is about ±9 points.
  • Item exposure in the population sample: the most used item appeared in 20.7% of sessions; 1,305 different items were used.

10. Norms

The current norms are provisional and model-based: IQ = 100 + 15θ, with θ on the scale fixed by the PROTEUS calibration of item difficulty. Three corrections named in the THETIS specification are built into the scoring module and can be switched on from the author's tools: a difficulty shift of the bank (Δb = 1.2), a linear link to Raven's Advanced Progressive Matrices (a = −0.15, b = 1.35) and a Pearson Type IV score distribution (skewness 0.5, kurtosis 4.0). They change scores substantially, so they stay off until norming data from human first sessions show which corrections the data support; the author's tools show their effect on sample scores before any change.

11. Data quality

Each first session carries four flags that never change a score: person fit (the standardized log-likelihood lz, Drasgow, Levine & Williams, 1985, judged against model-consistent sessions of the same ability), timing consistency, rapid wrong answers, and leaving the test page. Flagged sessions can be examined or set aside when norms are built.

12. Session delivery

Each answer is saved on the server as it is given, in Redis with a durable copy in Netlify Blobs; every write is a compare-and-set on a revision number, and each item carries a one-time token, so a retried or doubled request never counts an answer twice. The page retries every request with growing pauses, keeps an unsent answer in the browser, and resumes a reloaded or reopened session on the same item. A connection failure can therefore never change a score.

13. Limitations

  • Item parameters are a priori. Until they are recalibrated from human answers (after about 500 first sessions), the simulation shows the engine's precision under the model, not the test's validity.
  • Norms are provisional; above about IQ 160 scores rest increasingly on the item model, and above about IQ 190 a score is a position on the THETIS scale, not a population rarity.
  • No human validity data are published yet. Concordance with earlier proctored scores is being collected through the optional questionnaire.
  • The bank is thin from IQ 85 to 100 in four of the seven families.

14. Version history

  • Engine 2.0 (6 October 2026): untimed; four-parameter model with slip; routing block, mastery gate and crash protection; stop at SE 0.30 after 20 to 45 items; mandatory 14-item practice module; facet scores, confidence calibration and concordance in the report; Redis session store.
  • Earlier versions (October 2026): each item had a generous time limit; engine 1.0 also reported a processing-speed index; bank 1.1 made THETIS a power test scored on reasoning alone and added the Turnstiles and Orbits families; the test stopped at SE 0.28 after 18 to 34 items.

15. References

  • Barton, M. A., & Lord, F. M. (1981). An upper asymptote for the three-parameter logistic item-response model (RR-81-20). Educational Testing Service.
  • Drasgow, F., Levine, M. V., & Williams, E. A. (1985). Appropriateness measurement with polychotomous item response models and standardized indices. British Journal of Mathematical and Statistical Psychology, 38, 67–86.
  • Heinrich, J. (2004). A guide to the Pearson Type IV distribution (CDF/MEMO/STATISTICS/PUBLIC/6820). University of Pennsylvania.
  • Kingsbury, G. G., & Zara, A. R. (1989). Procedures for selecting items for computerized adaptive tests. Applied Measurement in Education, 2(4), 359–375.
  • Rulison, K. L., & Loken, E. (2009). I've fallen and I can't get up: Can high-ability students recover from early mistakes in CAT? Applied Psychological Measurement, 33(2), 83–101.