The paper makes lots of insightful observations, e.g., ceiling/floor issues, Working Memory deserving a careful examination, FRI/SAI not separating cleanly, etc…
There are few areas that require your attention though.
Probably the biggest overreach is your statement that the 4-factor model fits better than the 6-factor model. Because the 4-factor model is derived from EFA, it’s exploratory, thus, it’s a data-driven approach. What you have confirmed is that your model fits that particular sample well; if the sample used by the RIOT is less representative, this problem compounds. The ideal solution is to perform bootstrap and/or k-fold cross-validation, which you can’t do actually because you don’t have the raw data. There are 2 possibilities from here: admit the limitation with your approach or conduct a monte Carlo simulation, where you fit the 6 factor model, get the parameter estimates and treat this model as the true population model, then simulate datasets from that model, finally for each simulated datasets you’ll fit and compare both the 6-factor and your 4-factor models.
Actually, you may not even need to go that far. Because my view is that your paper already contained enough information that the 6-factor model is weak on a theoretical basis. Your strongest point is that some of the factors, namely, WMI, SAI, and FRI are hardly distinguishable. But look at what your Figure 5 produced: 2nd-order g loadings of .93, .98, .92. They are very close to 1. In some well known batteries, it happens sometimes that gf factor has a second order loading of 1 or close. But that’s just one factor. In the case where multiple, supposedly distinct factors have a relationship with g close to 1, it not only suggests that these factors are just g, but also that they are not distinct. Your 4-factor model, looking at figure 7, doesn’t have this problem. Interestingly, even the 5-factor model (figure 8) shows that FRI and VSI should not be treated as separate dimensions because their loadings are so close to 1. So I would say the 4-factor model is the best, but not for the main reason you invoked.
Unless you have time to conduct a Monte carle simulation, I think you should treat your 4-factor vs 6-factor model comparison as secondary evidence, while using Figure 5 (vs figure 7) as your primary evidence, because this is where the argument is strongest.
Relatedly, there are other details worth flagging, in your table 9. SRT is a Heywood case, and F6 is defined by CS subtest only, SRT loads much higher on RT factor compared to CRT. Those are big anomalies.
I don’t see whether the correlation matrix is already based on age-adjusted z-scores, or based on raw scores pooled across ages. It should be clarified because the difference matters!
On CFA assumptions, maybe you should state explicitly that, since you’re working with a published cor matrix, you can’t use robust tests to check for multivariate normality.
One caveat, regarding Table 15, you should acknowledge that the g-loadings may not be fully comparable, because Wechsler’s tests are normed on large, stratified, nationally representative samples, much larger than the RIOT’s N=417 (possibility not representative). Also, g-loadings will be affected if there are cross-loadings, but also by other factors such as test length, reliability (including floor/ceiling effects which impact reliability - which you have already mentioned were an issue in the test).
You are correct that CRT and SRT have equal loadings and this finding is anomalous, but I’d recommend you cite either Deary et al.’s paper “Reaction times and intelligence differences: A population-based cohort study” or Jensen’s “Clocking the Mind”, which support the idea that CRT is more g-loaded than SRT. More importantly, Jensen would argue against the idea of merging both. Jensen would in fact argue that SRT should be used as a covariate for studying the relationship between CRT and g (see Jensen 2006 pp. 70-71) because controlling for SRT raises the CRT-IQ relationship. Another issue I see with RTI is that SRT has an error rate near zero, and any error would not even be correlated with g (i.e., errors are not due to intelligence or lack of it). Thus, you need more than just SRT (which I consider as a warm-up test) and CRT to model the decision component and its relation with g, because decision is the component most relevant to psychometric g. Finally, browser-based timing is probably not reliable at the millisecond level. If you were to test people on reaction time, there’s no possible comparison between a real apparatus and online administration. Obviously, online tests are more convenient, but the cost is likely not benign, especially when it comes to measuring reaction time. This being said, my critique here may come off as a bit odd because the test is not supposed to measure reaction time the best it can, and by design it’s obviously not possible, but I think it should at least be informed.
For clarity, you should distinguish factor loadings from composite loadings, because in section 10 you noted “Model 2 structure, the VRI factor carries a solid .81 loading on g (Table 13)” and in table 13 the loading of .73 is displayed for VRI.
Regarding transparency, I can think of a few reasons why they want to restrict it, and one of these is misinterpretation by unqualified users. You don’t have to agree with this view, my point is that the creators can legitimately invoke that reason.
I accept the submission with 2 conditional changes: 1) do not rely on model fit indices for comparing 4-factor vs 6-factor model (at least, not as your primary argument) 2) make sure you clarify whether the correlation matrix used in Table 6 is age-adjusted or not.