Back to Submissions

1
An Independent Analysis of the RIOT IQ Test

Submission status
Reviewing

Submission Editor
Submission editor not assigned yet.

Author
Christopher G. Jacobs

Title
An Independent Analysis of the RIOT IQ Test

Abstract

The Reasoning and Intelligence Online Test (RIOT) reports a Full Scale IQ (FSIQ) score and six Cattell-Horn-Carroll (CHC) based index scores, but over a year since its release, its yet to be validated by an independant party. This study examines RIOT’s scoring, norming, reliability, and factor structure using publicly available data. Exploratory and confirmatory factor analyses were conducted on the 16 subtest intercorrelation matrix (N = 417), with RIOT's published hierarchical six-factor model used as a point of comparison. EFA best supported a four-factor model of Verbal Reasoning, Processing Speed, Reaction Time, and a merged factor containing the former Fluid Reasoning, Spatial Ability, and Working Memory subtets. CFA confirmed that a four-factor model inspired by the EFA provided the best fit (CFI = .966, RMSEA = .048, SRMR = .037) with one fewer free parameter than the published six-factor model. RIOT measures a moderately strong general cognitive factor, with FSIQ loading .86 on g in the best supported model. However, Fluid Reasoning and Spatial Ability did not separate cleanly, and the three Working Memory subtests failed to form a coherent factor. Overall, the evidence supports greater confidence in RIOT’s FSIQ, Verbal Reasoning, Processing Speed, and Reaction Time scores than in its Fluid Reasoning, Spatial Ability, and especially Working Memory indices. 

Keywords
intelligence, reliability, exploratory factor analysis, psychometrics, factor analysis, Confirmatory factor analysis, Cognitive assessment, RIOT IQ Test, IQ testing, Construct validity

Pdf

Paper

Reviewers ( 0 / 2 / 0 )
Reviewer 1: Considering / Revise
Reviewer 2: Considering / Revise

Wed 02 Sep 2026 21:15

Reviewer | Admin

It's a good study. I have no real complaints. I also checked some sessions for AI, though the text does not read like AI.

Regarding the choice between 4 vs. 5/6 factors. Note that g becomes essentially equivalent to fluid in the 5/6 models. This finding has often but not always been found in the literature, so it may be good or bad for those models. If you agree with the gf = g approach, where knowledge factors are essentially just a matter of information accumulation as a function of time and gf, then the identity of g and gf seen in the 5-6 factor models counts in their favor, even if the 4 factor model fits better empirically. Did you check the various models' modification indices to see if the poor fit is simply a matter of some missing covariances? This is often the case with complex models. I don't see any mention of modification index in the paper.

You could have a comparison table of sorts comparing the RIOT to plausible alternatives. A full WAIS test session may take about the same time, but would be much more expensive, especially if travel and wait time is taken into account. For another high quality online test, see CORE https://cognitivemetrics.com/test/CORE. As far as I know, there is no independent audit of that test either. Perhaps it could be your next project. Wonderlic is another commercial test, it is quite short (12 minute), and doesn't have all the group factor indices that many people like. Still, Wonderlic's g estimate may be quite close in quality to WAIS and RIOT.

Have you contacted RIOT/Warne to ask for data access? This could help with certain analyses. Or about the release date for the technical manual? Perhaps he would send you one privately to aid your study.

There has also been many factor analyses of other commercial tests. Their own preferred models are also often not found to be the best for their own data. Perhaps you could discuss this.

Reviewer | Admin

 

The paper makes lots of insightful observations, e.g., ceiling/floor issues, Working Memory deserving a careful examination, FRI/SAI not separating cleanly, etc…

 

There are few areas that require your attention though. 

 

Probably the biggest overreach is your statement that the 4-factor model fits better than the 6-factor model. Because the 4-factor model is derived from EFA, it’s exploratory, thus, it’s a data-driven approach. What you have confirmed is that your model fits that particular sample well; if the sample used by the RIOT is less representative, this problem compounds. The ideal solution is to perform bootstrap and/or k-fold cross-validation, which you can’t do actually because you don’t have the raw data. There are 2 possibilities from here: admit the limitation with your approach or conduct a monte Carlo simulation, where you fit the 6 factor model, get the parameter estimates and treat this model as the true population model, then simulate datasets from that model, finally for each simulated datasets you’ll fit and compare both the 6-factor and your 4-factor models.

 

Actually, you may not even need to go that far. Because my view is that your paper already contained enough information that the 6-factor model is weak on a theoretical basis. Your strongest point is that some of the factors, namely, WMI, SAI, and FRI are hardly distinguishable. But look at what your Figure 5 produced: 2nd-order g loadings of .93, .98, .92. They are very close to 1. In some well known batteries, it happens sometimes that gf factor has a second order loading of 1 or close. But that’s just one factor. In the case where multiple, supposedly distinct factors have a relationship with g close to 1, it not only suggests that these factors are just g, but also that they are not distinct. Your 4-factor model, looking at figure 7, doesn’t have this problem. Interestingly, even the 5-factor model (figure 8) shows that FRI and VSI should not be treated as separate dimensions because their loadings are so close to 1. So I would say the 4-factor model is the best, but not for the main reason you invoked.

 

Unless you have time to conduct a Monte carle simulation, I think you should treat your 4-factor vs 6-factor model comparison as secondary evidence, while using Figure 5 (vs figure 7) as your primary evidence, because this is where the argument is strongest.

 

Relatedly, there are other details worth flagging, in your table 9. SRT is a Heywood case, and F6 is defined by CS subtest only, SRT loads much higher on RT factor compared to CRT. Those are big anomalies.

 

I don’t see whether the correlation matrix is already based on age-adjusted z-scores, or based on raw scores pooled across ages. It should be clarified because the difference matters!

 

On CFA assumptions, maybe you should state explicitly that, since you’re working with a published cor matrix, you can’t use robust tests to check for multivariate normality.

 

One caveat, regarding Table 15, you should acknowledge that the g-loadings may not be fully comparable, because Wechsler’s tests are normed on large, stratified, nationally representative samples, much larger than the RIOT’s N=417 (possibility not representative). Also, g-loadings will be affected if there are cross-loadings, but also by other factors such as test length, reliability (including floor/ceiling effects which impact reliability - which you have already mentioned were an issue in the test).

 

You are correct that CRT and SRT have equal loadings and this finding is anomalous, but I’d recommend you cite either Deary et al.’s paper “Reaction times and intelligence differences: A population-based cohort study” or Jensen’s “Clocking the Mind”, which support the idea that CRT is more g-loaded than SRT. More importantly, Jensen would argue against the idea of merging both. Jensen would in fact argue that SRT should be used as a covariate for studying the relationship between CRT and g (see Jensen 2006 pp. 70-71) because controlling for SRT raises the CRT-IQ relationship. Another issue I see with RTI is that SRT has an error rate near zero, and any error would not even be correlated with g (i.e., errors are not due to intelligence or lack of it). Thus, you need more than just SRT (which I consider as a warm-up test) and CRT to model the decision component and its relation with g, because decision is the component most relevant to psychometric g. Finally, browser-based timing is probably not reliable at the millisecond level. If you were to test people on reaction time, there’s no possible comparison between a real apparatus and online administration. Obviously, online tests are more convenient, but the cost is likely not benign, especially when it comes to measuring reaction time. This being said, my critique here may come off as a bit odd because the test is not supposed to measure reaction time the best it can, and by design it’s obviously not possible, but I think it should at least be informed.

 

For clarity, you should distinguish factor loadings from composite loadings, because in section 10 you noted “Model 2 structure, the VRI factor carries a solid .81 loading on g (Table 13)” and in table 13 the loading of .73 is displayed for VRI.

 

 

Regarding transparency, I can think of a few reasons why they want to restrict it, and one of these is misinterpretation by unqualified users. You don’t have to agree with this view, my point is that the creators can legitimately invoke that reason.

 

I accept the submission with 2 conditional changes: 1) do not rely on model fit indices for comparing 4-factor vs 6-factor model (at least, not as your primary argument) 2) make sure you clarify whether the correlation matrix used in Table 6 is age-adjusted or not.