scieee AI-readable full text Open interactive document viewer

Learning Cyber Security or Only Memorizing?

Koskinen, J. A.

Abstract

A self-study course was designed in 2020 in such a way that the explicit objective of study would be the exams based on multiple-choice questions. The idea was to let students do what teachers often complain they do anyway, and to infiltrate some learning as a byproduct of the memorizing procedure. Investigating how the questions are solved shows that this approach is suboptimal. Partial question disclosures, intentional or not, and practice exams are common, whence the result and how it was obtained can be useful elsewhere too.

Full text

Practice Paper Recommended citation: Koskinen, J. A. (2025). Learning Cyber Security or Only Memorizing?. In Kangaslampi, R., Langie, G., Järvinen, H.-M., & Nagy, B. (Eds.), SEFI 53rd Annual Conference. European Society for Engineering Education (SEFI), Tampere, Finland. DOI: 10.5281/zenodo.17631832. This Conference Paper is brought to you for open access by the 53rd Annual Conference of the European Society for Engineering Education (SEFI) at Tampere University in Tampere, Finland. This work is licensed under a Creative Commons Attribution-NonCommercial-Share Alike 4.0 International License. LEARNING CYBER SECURITY OR ONLY MEMORIZING? J. A. Koskinen 1 Tampere University, Finland, ORCID: 0009-0004-9019-2217 Conference Key Areas: Open and online education, Digital tools Keywords: exam-oriented course, multiple-choice question, cyber security ABSTRACT A self-study course was designed in 2020 in such a way that the explicit objective of study would be the exams based on multiple-choice questions. The idea was to let students do what teachers often complain they do anyway, and to infiltrate some learning as a byproduct of the memorizing procedure. Investigating how the questions are solved shows that this approach is suboptimal. Partial question disclosures, intentional or not, and practice exams are common, whence the result and how it was obtained can be useful elsewhere too. 1 INTRODUCTION Since autumn 2020 we have been running a self-study course heavily based on a pool of multiple choice-questions (MCQs) in a basic course of cyber security. Students have taken supervised exams at their own pace and gathered extra points for their grade from a set of unsupervised exercises. The collection of 645 MCQs strives to cover the course objectives in a definitive manner. Only now, during the 5th year of the course, have most MCQs been used sufficiently many times to enable quantitative analysis of their quality. This research started as an attempt to use item analysis to justify and share our best ideas for creating MCQs in this field. Most topics in cyber security were difficult to represent adequately with MCQs. This may seem odd, while many readers have probably encountered security awareness tests or training, where MCQs appear. Their answers often depend more on organizational rules than cyber security principles. Policies and rules must be understood, of course, but we wanted students to learn ideas from MCQs and apply them to solve the same and varied MCQs in their exams. We have failed. Students memorize the MCQs and not what they mean, or at least that is how they behave in their exams. This observation may be significant e.g. for MCQ-based certifications, like those from (CISCO, CompTIA, ISACA (n.d.)). 1 [email protected] We first describe the course, its exercises, and the exams, because of details that support our conclusion. Then we explain with five examples how we cannot but draw that conclusion. 1 Characterization of good cyber security MCQs is left to some other time (meanwhile, check Newton, 2023 and Liu &al, 2023). We assume the reader is familiar with the ordinary exam item analysis (see e.g. Reynolds &al, 2021). 2 CONTEXT 2.1 The course The following table describes the course components on the right and associates the parts of the MCQ pool with them on the left. The course is continuously available without any schedule, but the exams form a strict succession. Each MCQ has one correct option and three distractors. Scoring is with 1, 0, and −1/3 points for correct, empty and wrong answer. Both Exam 1 and 2 have a passing limit of 22 points, and 12 is subtracted from the score to calculate the course grade. MCQ pool Component consists of brings points Grade obtainable public hidden total MCQs pages 141 57 198 Exam 1 32 33 10..20 1 300 147 447 Exam 2 52 +comments 84 10..40 1 .. 3 441 204 645 Exam 3 essays 10..20 1 .. 5 Exercises: 25 automatic, 5 specials 2 0..25 (out of 38) The MCQs were created by translating into English the most suitable ones from earlier courses in Finnish language, and crafting 251 new ones. All were written by the author of this paper and proofread by two colleagues. The 141 public Exam 1 questions are given on 33 pages that describe very briefly the core topics. Each page has a few hidden questions, and the 32 questions are drawn so that 7 questions are from the hidden ones and no two questions are from the same page. Exam 2 has the same principle, but its 52 questions come from 84 pages and 11 questions are from the hidden pool. The Exam 2 pages have much more content than the Exam 1 pages. While being relevant to the topic, that content is not directly intended to solve the related MCQs. 2.2 Exercises Automatic exercises are based on multiple-choice or short-answer questions. Nine exercises require some activity like computation or using software. Students can choose which exercises they do and proceed toward the maximal 25 points regardless of failures or partial points, because there are 38 points on offer in total. All the exercises are the work of the current author except one with MCQs following the CISSP ontology and being adapted from publicly available practice questions. 3 Most of them were far too transparent or they dealt with very specific 1 Our direct check by emailing 140 students received one response. The student, who achieved grade 5, wrote that he worked by first understanding but then collected the pool and memorized it. 2 The specials are: contributing to the materials, reporting news, interviewing, organizing one’s storage, and writing essays. The first, and partly the second, represent flipped learning (see e.g. Baig & Yadegaridehkordi, 2023) and result in creation of course material that is otherwise (intentionally) missing. This is a link showing the contributions, and also the public MCQs: tie-sec.rd.tuni.fi/maso/. 3 CISSP, Certified Information Systems Security Professional is a popular skill certificate that is based on a 6-hour exam of 250 MCQs. Practice exams can be seen e.g. at https://www.cissp.com. terms that are not in line with the purpose of our course. The basic idea of our MCQdriven course is condensed in the instructions of this particular exercise, containing: “Hopefully you use your unlimited attempts by learning what distinguishes the right answer from the wrong ones, and not just by using brute force to pass. If you concentrate on the matters and find out more information, you will also have a safe way to practice how to learn from the MCQs that are on display for Exam 1 and 2.” On average students scored 0,85 points (out of 1) from each of the automatic exercises, and the proportion of exercise points out of the total grade points is averagely 23%, being naturally largest (44%) for grade 1, and increasing from 16% for grade 2 and 3 to 20% for grade 5. There is great variation between the students, as summarized by the correlation −0,2 between exercise points and exam points. 2.3 The exams, with results Besides appearing on the topic pages, the public MCQs can also be seen in practice exams, one available before each attempt of a real exam. The following table shows the number of exams taken, the number of students involved, and their average results. The number of practice exams taken was 169 and 102. Exam Exams taken Students taking, passing Passers’ average score, time per MCQ 1 217 152 140 27,3 40’’ 2 136 110 106 40,3 1’6’’ Exam 2 went somewhat better than Exam 1, although it is slightly more advanced. The averages lie at 53% and 61% of their respective ranges (22..32 and 22..52). There are several reasons behind this. Most obviously, Exam 1 can be attempted as many times as needed but passed only once, while Exam 2 has three attempts and improving a passing score by one or two retakes is possible. Secondly, it is likely that those who proceed to Exam 2 not only have higher goals – which may also include Exam 3 – but also work harder for them than those who confine to Exam 1 and pass the course with the lowest grade. This is probably evidenced by the 65% more time they spent on each MCQ. Furthermore, a small effect comes from the commenting feature that Exam 2 has. It was originally designed to aid the teacher with immediate quality control of the questions, later augmented to give more control for the students into the MCQ situation which is commonly considered somewhat unnatural. Commenting works in the following way. During Exam 2, after seeing their results, students can write comments to any questions. Giving truly good excuses for two wrong answers might get their negative impact neutralized, in principle even the positive point given by the teacher. Only a handful of neutralizations happened and no reversions. With the augmented feature, two points can be obtained by showing good understanding of the meaning of four correctly answered MCQs. Altogether there were 222 comments written by 52 students, who scored averagely 0,74 extra points. While 21 students scored 0, those others who had understood the commenting feature better obtained averagely 1,24 extra points. In any case it is surprising why so few, less than half, even tried to use this feature. If it is true that the students memorize the MCQ answers this becomes self-evident. Note that this is related to Exam 2 regardless of the higher effort given to it. 3 GENERAL STATISTICS OF THE MCQS 3.1 Modified scoring and averages The MCQ-wise calculations that follow need not use the scoring with 1, 0 or −1/3 points. Instead, we make the correct answer slightly more balanced by first replacing those scores with integers 3, 0 and −2 (instead of 3,0,−1). Then we fine-tune a little by considering the corrections students made in their exam. The correction feature in both Exams 1 and 2, real and practice, works like this: After submitting their answers the students see their total score but not which MCQs were wrong or right. At that stage they can switch any answer they want, including making an answer empty. Only after this stage, the student concludes the exam and sees their MCQ-wise results. The total score is computed on the final choices – after any corrections. The following table explains the fine-tuning and scoring and gives the distribution of points among the 24 522 MCQ instances in both real and practice versions of Exams 1 and 2. From the very small proportion (4,34%) of other than the 3, 0, −2 scores we infer that it would not be useful to do such fine-tuning when grading the students. For the purpose of our analysis, the effect is also going to be small. The next table gives averages and counts for various kinds of MCQs. Public questions Hidden questions Exam 1 Exam 2 Exam 1 Exam 2 Both Exam 1 Exam 2 Both Real avg. 2,03 2,01 2,27 2,38 2,33 1,17 0,61 0,89 N 6848 7052 5384 5566 10950 1496 1486 2982 Practice avg. 1,66 1,54 N 5332 5290 The averages show no big difference between questions of Exams 1 and 2, overall or on the public side. There is a big difference, nearly double, in the hidden side. This means that the better performance in Exam 2, as elaborated in section 2.3, may not liberate it from our memorizing hypothesis. For the practice exams the averages here show that students are clearly less prepared for them. There were probably many practice sessions where the student used some external aid – not mimicking the real situation. Hence, on average the students were indeed just practicing. 3.2 Distribution among topic domains The final table of general statistics, on the next page, shows how the MCQs were distributed among the 11 domains of our course and where difficult MCQs occurred. Difficulty is defined here as average score being below −1. Limit for easiness is 2,9 points. The three other columns show the proportion of questions that had correlation in the indicated ranges (“good”, “poor”, “bad”). The two ‘N’ columns show absolute quantities of MCQs out of which the proportions were calculated. The difference between public and hidden MCQs was expected and can be clearly seen. But it was surprising to see how big a part of MCQs correlate negatively, especially the really high proportion in the hidden MCQs of the two first domains. Our criteria for difficult and easy may be quite strict. We will elaborate on some examples of difficult questions later. For easy questions we just note that our course needs them for coverage because the public MCQs are supposed to define the learning objectives. Easy questions are also important to motivate students in a points count % fine-tuning based on corrections −3 165 0,66 wrong to wrong −2 3 486 13,94 unchanged wrong, or empty to wrong −1 201 0,80 correct to wrong, or wrong to empty 0 2 770 11,07 unchanged empty 1 227 0,91 wrong to empty 2 492 1,97 correct to empty 3 17 181 68,68 unchanged correct, or empty to correct couple of ways: They are more likely to see why one option sticks out among the distractors, which may lead to understanding the “message” of the question. Easy questions may encourage students to tackle more difficult ones in the same manner during the learning phase, and during the exam they may keep up the hope that the task is doable. 4 COMPARISONS INDICATING MEMORIZATION There are many sets, usually pairs, of variations where one MCQ is public and the other hidden. We picked five sets to show how much students seemingly memorize. Our results can be contrasted to those of Little & Bjork (2015) suggesting that earlier plausible distractors can lead to improved results later, and of Butler (2018) suggesting that MCQs can promote learning by repeated exposure to similar questions. 4.1 Too sour Acid? (Exam 2) The same 1-word answer options Atomicity, Consistency, Isolation, and Durability appear in three MCQ variants starting with: “In database management systems, ACID is an important set of properties. Which one is to allow …”. The following table shows the continuations of the stems and the results in the exam (frequencies and score correlations). Correct answers are indicated in bold green. A C I D corr. public …only valid transactions that do not violate any user-defined integrity constraints? 2 9 4 1 0,32 …no partial transactions to be visible before the whole transaction is completed? 2 1 15 1 0,63 hidden … no partial transactions to happen if the whole transaction is not completed? 0 2 7 0 -- In the hidden variant the correct option Atomicity was not picked by anyone. Most popular was Isolation as if memorized through the fairly similar-looking stem in the public variant. The common meaning of isolation is quite far from what the stem expresses. On the contrary, there is a possibility that Consistency is misinterpreted to correspond to what the hidden stem says. On the other hand, Consistency could be memorized from public side by those who realized that Isolation is not good here. Although the small number of occurrences gives less significance to the quantities and correlations, the public–hidden difference shows a high extent of memorization – especially because it seems unreasonable to judge the hidden variant as too difficult. Apparently it would still be good to modify it to attract more correct answers. 4.2 Was the shared secret kept in mind? (Exam 1) Our second example is a pair of MCQs with identical stem: “Shared secret is an InfoSec term that usually refers to…”. The correct answer in the public version is “…a result of protocols like IKE (Internet Key Exchange) of IPsec.” and it received 30 out of the 39 responses (one being empty). The correlation was also quite high, 0,63. The hidden version has this correct answer: “…two parties but no one else knowing the same bit sequence that allows them to authenticate each other.” It is semantically very close to the public case and also close to the entire definition of shared secrets – mentioning confidential messaging would complete it. The most popular option in the hidden version was: “… the bit sequence that protocols like IKE (Internet Key Exchange) of IPsec fetch from a trusted server.” This received 16 responses in contrast to the 2 responses for the correct option. Additionally, one response was empty. The support for the memorization hypothesis lies in the possibility that the sophisticated term IKE was kept in mind from the public version. Not all are memorizing however – or successful in it. You may skip the technical discussion of this. 1 4.3 What is certification? (Exam 2) The public–hidden pairs need not share the stem to become indicators of memorization. For instance, the public version with stem “An X.509 certificate is used to verify…” gets all 17 responses at the correct option which is very close to the definition of a certificate: “…the binding between an identity and a cryptographic key.” The hidden version has a different stem: “An X.509 certificate can be obtained…” but the majority (7) of responses concentrate on a similar-looking but wrong binding idea: “…to prove the binding between a private key and the corresponding public key.” If this were the result of the students’ actually applying their learning it would be unfortunate. Firstly, the binding between the private and public key is mathematical and intrinsic to the cryptosystem itself. Secondly, the correct option “…to assure the web browsing users of the identity of a website that uses TLS.” is at the heart of internet security and the main use of certificates. Only one student picked this option. It is understandable that trying to pass the exam quickly and having an opportunity for retakes may cause sloppy reading of questions. Two distractors in the hidden version indicate something like that: Distractor “… if the information system has been audited according to the CERT standard.” was picked by one, and “…to prove that an information system has passed certification against any one of the criteria models.” by two students. Both options refer to a different kind of certification than X.509. 2 1 These two distractors were picked by some: “… personal sensitive data that cannot be only private but must be known to someone else, too, than the owner.” (3 students), “… the random bit sequences that two parties send to a trusted server for the purpose of authenticating each other.” (5 students). The first one would be quite good if the concepts of personal and sensitive were not reserved for matters related to privacy in the sense close to intimacy. The second and more popular distractor similarly includes something correct (like the purpose) but actually it is further from truth than the first distractor. Both distractors can be seen as successful by the usual criteria of MCQ construction. While the hidden MCQ is too difficult even with our strict limits (averaging −1,4 points), it is doubtful whether textual completion of the correct option will make it easier in the current situation. 2 …and that certification is different from the skill certification mentioned in Introduction. 4.4 Is the asymmetry of a cryptosystem too complex? (Exam 1) Further evidence of transfer from an already seen option to a hidden option is shown in this example, where the topic conveniently provides a 2x2 grid for options yielding four variants: An asymmetric cryptosystem consists of a key pair, public and private. A public key can have two uses: encryption and verification, while a private key can be used for decryption and signing. For easier comparison we abbreviate: CAPK = cryptographic analogue of a physical key, PoB = padlock on a box. (Note, each of MCQ #1 and #3 also received one empty response.) The arrows do not indicate observations between the answers of any individual student. We use them to display a possible memorizing influence from the publicly available MCQs toward the hidden variants. The reasoning again goes technical 1 . Besides in the materials, students may have seen questions #1 and #2 in their practice exam. We investigated whether the memory influence could be partially explained that way. There were 26 instances for each of #1 and #2 in practice exams where they appeared right before any of the four questions (#1, #2, #3, #4) in a real exam. Practice #1 was clearly easier than #2 (averages 2,0 and 0,85). All 7 who had #1 repeating, got it right, while 3 out of 9 failed with a repeating #2. The most difficult question, #3, went only slightly better in the real exam (average −0,29) for the 14 who had seen either #1 or #2 in the practice exam, where their average was a reasonable 1,9. There were 36 students having a different question in their real and practice exam. Their scores had a negligible correlation −0,08. (The already treated situation of a repeated question had correlation 0,37 for 16 students). These observations match well with the idea that the ability to answer correctly in the real exam is dependent on memorizing questions somewhat blindly. Question #3 was not only the most difficult among these four, but it ranked 11 in difficulty among all 645 MCQs, and it was the most difficult one among those that 1 The public questions were solved very well with two possible answers encrypt or verify, (receiving 84% and 81%). To distinguish them, possibly rules “shut→encrypt” and “open→verify” were learned. The hidden variants received a large portion of wrong answers with those rules: 11 (28%) for “shut” and 11 (42%) for “open”. Even without a memory rule, the most advanced concept, signature verification, after seeing it in #2 may have influenced the high proportion of picks (15; 38%) in question #3 (dotted line). Signature creation, the correct answer of #3, is one step simpler conceptually than verification, but still it received only 7 picks, and also fewer picks than encryption. Neither encryption nor verification is what private keys are for. had at least 20 appearances. For the sake of students’ progress, it might be better to discard it or modify it with some hints. However, it still had a positive correlation (0,27), and the symmetry in this setting of four questions makes us think that it is worth preserving – maybe by switching its publicity with the easiest question #1. 4.5 What did a beacon illuminate? (Exam 2) In this example the MCQs share the same stem and one option, and the outcome is still very different. The public one was very easy but correlated nicely (0,73): “A web beacon, or web bug…”. with the correct option receiving 20 picks: “… is a web page element that the user is not supposed to notice.”. Two of the distractors received altogether 3 picks, and this distractor none at all: “…is an invisible image loaded within a web page that cannot be associated with the use of cookies.” However, this was a distractor also in the hidden variant where it received 8 picks, while the correct option “…comes to the web page usually from a different server than the main content.” only received 1. In addition, there were 3 empty answers. Is this also a memorization issue: the correct option in the public version is in the form of a definition, and so is the most popular option in the hidden version – unlike the correct option. Albeit with small quantities, this case might indicate that trying to teach with MCQs has failed also in a different way: Students who had studied the public version had not taken note of the wrongness of its distractors. When they didn’t see the memorized correct option in the hidden version they picked the option that mostly resembled what they had seen. Of course, the meaning of the first half of that option (invisibility) is correct, while the rest is not. 5 CONCLUSION AND WAY FORWARD In a way students behave as we expect – learning from MCQs – but our expectation is not fully realized and has led to suboptimal performance in exams, possibly also to superficial learning. However, our inference is indirect and this report is based on a limited number of worst cases, albeit with some supporting evidence. Of course, most students also learned in a real sense, but still, in the long run the results of this investigation may lead us to abandon the current structure and try something more traditional – or new, like organizing the public MCQs into a game, together with the exercises. Before that we should carry out a study involving a preand post-test of students’ skills with an instrument that does not involve the current set of MCQs. It would also be interesting to see the first scientific study on the “memory effects” on some of the popular commercial certifications. Memorizing may not be the strategy students use intentionally when learning but it may take over during the exam through intuition or recognition of familiar patterns. Augmenting assessment with a new dimension would let us distinguish if such intuition is at play. For research, a simple approach would be to ask at every MCQ, whether the student is certain of having seen exactly that question. For teaching, this could be accompanied by asking whether the student is certain about their answer – and halving points for uncertain answers. In the short run, the observations reported here lead us to improve the most difficult questions. Besides changes in their wording and concepts, one straightforward way to save them – as such – is to move them to the public pool, which lets students “learn” them in advance. For many of them a simplified version ought to be created into the hidden pool, to maintain its size and variety. Increasing the size of the public pool makes pure memorizing more difficult and may promote learning in the desired way.