The abysmally low score in Experiment 1 suggested to test the sensitivity of the scoring algorithm with respect to:
(1) answering incorrectly problems in the begining of the test when it is still oscillating wildly in problem difficulty, trying to adapt to the test-taker level
(2) answering incorrectly several questions in a row, thus rejecting the problem difficulty oscillations and biasing the algorithm towards offering later problems of lower difficulty.
It was said on this forum these were problems in the 90's with the GRE adaptive algorithm which was later abandoned. Unfortunately the following experiment, still shows those problems with the current GMATPrep algorithm, which doesn't surprise me at all.
Experiment 4:
--------------
- used "Test 1" in GMATPrep
- answered questions 1,2,3 and 6,7,8 purposely incorrectly
- answered all other questions correctly
Results:
- questions 37 total: incorrect 6 (the intended questions 1,2,3,6,7,8), correct 31
- scaled score: 47 corresponding to 76%
Notes:
I got only one very hard question (one of the very hard questions in Experiment 3). All other questions were either average difficulty or hillariously easy - easier than in Experiment 3. I have all questions recorded so I can show anyone what I mean.
The level after question 8 had dropped so dramatically that in question 9 I was asked if (x^6)(x^4) = x^10. That was expected after answering questions 1,2,3 and 6,7,8 wrong. What I found disappoing was that the scoring algorithm never recovered to offering higher level problems, it kept giving me easy problems, despite the fact I solved the next 29 questions correctly in a row without a single error.
The final result reflects that algorithmic flaw. I got only 6 questions out of 37 wrong, the algorithm thinks my true level is in the 76th percentile. That means that according to GMATPrep, every 4th person does better than that - better than getting 31 out of 37 questions correctly.
It seems that the algorithm decides what the test taker level is from the answers of the first questions, and for the later problems does not attempt wild oscillations in problem difficulty as it did in the begining of the test, it just keeps giving problems around the same old level.
The problem with that is that the level is decided in the first questions where most test takers make the silliest errors before they have warmed up enough. The second problem is that the algorithm doesn't give the test taker a second chance by drifting towards higher levels later if the test taker keeps answering correctly: I was delegated to the 76th percentile just because I had answered 6 of the first questions incorrectly.
Experiment 4 clearly suggests the adaptive scoring of GMATPrep is far from ideal and suffers from the same problems of the GRE adaptive scoring in the past. This may or may not translate to the real test scoring because the real test has a significantly larger bank of problems and may react more adequately to test takers that underperform in the begining of the test.
A good experiment for the future would be to answer questions 1, 2, and 3 purposely incorrectly and see what happens later when the test-taker answers all the other questions correctly.