scieee AI-readable full text Open interactive document viewer

Building symbolic representations of intuitive real-time skills from performance data

D. Michie,R. Camacho

Abstract

Real-time control skills are ordinarily tacit - their possessors cannot explicitly communicate them. But given sucient sampling of a trained expert's input-output behaviour, machine learning programs have been found capable of constructing rules which, when run as programs, deliver behaviours similar to those of the original exemplars. These 'clones' are in efect symbolic representations of subcognitive behaviours. After validation on simple pole-balancing tasks, the principles have been successfully generalized in flight-simulator experiments, both by Sammut and others at UNSW, and by Camacho at the Turing Institute. A flight plan switches control through a sequence of logically concurrent sets of reactive behaviours. Each set can be thought of as a committee of subpilots who are respectively specialized for rudder, elevators, rollers, thrust, etc. The chairman (the flight plan) knows only the mission sequence, and how to recognize the onset of each stage. This treatment is essentially that of the 'blackboard model', augmented by machine learning to extract subpilot behaviours (seventy-two behaviours in Camacho's auto-pilot for a simulated F-16 combat plane). A 'clean-up' effect, first noted in the pole-balancing phase of this enquiry, results in auto-pilots which fly the F-16 under tighter control than the human from whom the behavioural records were sampled.

Full text

15 BuildingSymbolicRepresentations ofIntuitive Real-time Skills from Performance Data D. Michie and R. Camacho The Turing Institute, Glasgow, UK Abstract Real-time control skills are ordinarily tacit | their p ossessors cannot explicitly communicate them. But given sucient sampling of a trained exp ert's input{output b ehaviour, machine learning programs have b een found capable of constructing rules which, when run as programs, deliver b ehaviours similar to those of the original exemplars. These `clones' are in eect symbolic representations of sub cognitive b ehaviours. After validation on simple p ole-balancing tasks, the principles have b een successfully generalized in ight-simulator experiments, b oth by Sammut and others at UNSW, and by Camacho at the Turing Institute. A ight plan switches control through a sequence of logically concurrent sets of reactive b ehaviours. Each set can be thought of as a committee of subpilots who are resp ectively sp ecialized for rudder, elevators, rollers, thrust, etc. The chairman (the ight plan) knows only the mission sequence, and how to recognize the onset of each stage. This treatment is essentially that of the `blackboard mo del', augmented by machine learning to extract subpilot b ehaviours (seventy-two b ehaviours in Camacho's auto-pilot for a simulated F-16 combat plane). A `clean-up' eect, rst noted in the p olebalancing phase of this enquiry, results in auto-pilots which y the F-16 under tighter control than the human from whom the b ehavioural records were sampled. 385 REAL-TIME SKILLS Table 15.1. Criteria of strong and weak AI Strong Weak Feasibility Human-level intelligence Human-level intelligence of goals will b e achieved in will b e implemented only machines within in some unimaginable foreseeable time. future, or p erhaps never. Forms of All thought can be Most thought is implemenmechanized as sequential intuitive, not tation logical reasoning from introsp ectable, axiomatic descriptions non-logical, asso ciative, of the world. The approximate and `physical symbol system `fuzzy': b est mo delled hyp othesis': all agents, by brain-like including intelligent, ultra-parallel networks. are b est implemented symb olically. Personnel Vintage AI professionals, Memb ers of other e.g. Turing, Simon, professions, particularly Newell, McCarthy, in linguistics, Feigenbaum, Nilsson, and neurobiology, physics, their followers. and philosophy. 1. INTRODUCTION The lab els `strong AI' and `weak AI' have sometimes b een used to dierentiate two scho ols. Criteria are summarized in Table 15.1. The taxonomy in Table 15.1 lays emphasis on the `physical symb ol system hyp othesis' of Newell and Simon (1976). Their intended interpretation restricts symb ol systems to those which can transparently supp ort communication with human users. Thus the lists of numerical weights in which neural nets express themselves constitute `symbols' of a sort, but not in the sense intended by the above authors. This restriction has p ersuaded some practitioners that the physical symb ol system hyp othe386 D. MICHIE AND R. CAMACHO sis excludes intuitive pro cesses from AI's domain of discourse. Such separatism is unsafe, since much knowledge-based thought seems irredeemably intuitive and sub-articulate (for a recent commentary see French, 1990). For its sub cognitive pro cesses, there is no direct evidence that the brain employs a symbolic regime. Hence those who accept subarticulate exp ertise as a prop er AI concern may wonder whether for this purp ose they should abandon symb olic representations as untrue to nature. The present chapter advo cates a dierent p osition, namely that a conceptually transparent symb olic style oers a way of improving on nature. By representing intuitive pro cesses symbolically, inductive inference can do something which is b oth non-brainlike and also highly useful, catering to the client who says: `My inhouse exp erts may b e `intuitive'. But I want an exp ert system to formulate its reasons more explicitly than that.' 2. KNOWLEDGE AND THOUGHT In industrial knowledge systems the implementer has to distinguish b etween thought as something to b e communicated and thought as problem solving. Choice of representation remains a develop er's option. In implementing intuition, he or she may decide that it is something over which to draw a veil. The veil may b e woven of neural nets, or of hand-crafted spaghetti-co de, or of something else. But suppose that the develop er has to supply the customer also with means to draw the veil aside, for purp oses of interrogation ab out goals, plans, evidence, justi- cation, and the like. At the price of b eing less true to nature, he or she might then b e b etter o not to have veiled it in the rst place. Like cognitive and brain scientists, knowledge engineers also study the structure of exp ertise. Unlike cognitive and brain scientists, they do this (or should do) for the purpose not of emulating but of transcending the brain's limitations. First among these is the relative inarticulacy of what b oth cognitive scientists and knowledge engineers call `pro cedural knowledge', thus distinguishing it from `declarative'. 387 REAL-TIME SKILLS 2.1. Declarative knowledge It is characteristic of the retrieval and use of declarative knowledge that it is ordinarily done in conscious awareness. From a wealth of neurobiological observations concerning the eects of brain lesions on memory, L. R. Squire (1987, chapter 11) distinguishes declarative memory from pro cedural as `memory that is directly accessible to conscious recollection'. By contrast, the hall-mark of a highly trained exp ert brain is that it do es much of its work intuitively. `Dialogue elicitation' of rules for building exp ert systems may therefore b e frustrated whenever a given exp ertise involves strategies stored in pro cedural memory. Inaccessibility to consciousness of even parts of a targeted exp ertise can then cause serious problems for large knowledge engineering pro jects, such as Japan's ambitious `Fifth Generation' (Michie, 1988). Dierentiation of the two forms is thus desirable. Declarative knowledge comprises whatever lends itself to logical formulation: goals, descriptions, constraints, p ossibilities, hyp otheses. The declarative category also includes facts. When these relate directly or indirectly to events in the agent's own exp erience, their place of storage is referred to as `episo dic' memory. Another sub division of declarative knowledge is held to reside in `semantic' memory, which Squire denes as follows: Semantic memory refers to knowledge of the world. This system represents organised information such as facts, concepts, and vo cabulary. The content of semantic memory is explicitly known and available for recall. Unlike episo dic memory, however, semantic memory has no necessary temp oral landmarks. It do es not refer to particular events in a person's past. A simple illustration of this dierence is that one may recall the dierence b etween episo dic and semantic memory, or one may recall the encounter when the dierence was rst explained. A scho ol founded by John McCarthy (1959) aims to extend formal logic to serve as a vehicle for mechanizing declarative knowledge (see a recent collection edited by Ginsb erg, 1987). We will say little further ab out the pro ject, b eyond expressing resp ect for such work. Its philosophical imp ortance is matched 388 D. MICHIE AND R. CAMACHO only by its diculty. Our theme is closer to the name and nature of exp ert systems. These are not so much to do with giving computers knowledge of the world, as with equipping them with useful know-how. In face of the diculties which confront the McCarthy pro ject, there is something to b e said for separately studying the mechanization of pro cedural knowledge and only later integrating the two levels. 2.2. Nature of procedural knowledge In Anderson's (1990) text on cognition, skilled pro cedures are pictured as arising in part by derivation from pre-existing mental descriptions. No direct evidence is oered. Knowledge engineers concerned with real-time skills have b een led by practical exp erience in a rather dierent direction. The empirical picture is one of inductive compilation from sensorimotor data gathered in the course of trial and error. In this picture the role of higherlevel knowledge is not to participate directly, but to steer the learning pro cess, setting and adjusting the frame within which skill-b earing rules are constructed. The nal phase of skill-learning, describ ed by Anderson and others as `automatization', do es not ordinarily supp ort introsp ective rep ort by the exp ert p erformer, hence the `knowledgeacquisition b ottleneck' of applied AI. Procedural knowledge, as we have seen, limits itself to the `how to' of skilled tasks, whether physical as in making a chair, or more abstract as in prediction of sterling rates against the dollar or the diagnosis of acute ab dominal pain. A common synonym for such knowledge is know-how, and its manifestation in observable b ehaviour is called `skill'. One diculty is that observed task-p erformance do es not necessarily reveal whether a given expert's b ehaviour really exemplies a skill in the pro cedural sense or whether he or she is using declarative-semantic memory to form action-plans on the y. Squire's earlier-cited denition supplies a test, namely the ability to give a verbal account of the way in which each decision was made, p ossible only for declarative memory. A second criterion is the frequency of the recognize{act cycle: this may simply b e to o fast for `what-if ' inferential planning to b e feasible. For those concerned to recover pro cedural rules, as in build389 REAL-TIME SKILLS ing exp ert systems, lack of verbal access (on which Anderson also remarks) is a problem. Yet there is widespread faith among knowledge engineers that special methods of `dialogue elicitation' can be found which will p ermit the construction of rulebased systems on the scale of such inductively built systems as the GASOIL (Slo comb e et al. , 1986) and BMT programs of Table 15.2. Is rule induction from expert-supplied data nevertheless in some sense a second-b est option for building systems on the BMT scale? On the contrary. Exp erts can rapidly and eectively communicate their skills (as in the BMT case) solely via illustrative resp onses to selected cases. Do es he or she thereby omit something indispensable? Certainly the practitioner's explicit and communicable awareness is basic to exp ertise in some task domains. But other domains, which lack this prop erty, can b e found not only among a rather wide variety of industrial tasks, but even in such purely `mental' forms of exp ertise as playing a strong game of checkers (see b elow). As a paradigm of pro cedural knowledge, Feigenbaum and McCorduck (1983, p.55) give the example of tying one's sho es. It is interesting that once this skill has reached the stage known as automatization it can continue unaected by destruction of the individual's brain mechanisms for acquiring and handling imp ortant forms of declarative knowledge. Damasio describ es a patient named Boswell. The following summary is from Patricia Smith Churchland (p ersonal communication). In addition to losing the hipp o campal structures, he has massive damage to frontal cortex. He can identify a house, or a car, but he cannot identify his house or his car; he cannot remember that he was married, that he has children, and so forth. He seems to have no retrograde episo dic memory, as well as no anterograde episodic, ... Boswell can still play a ne game of checkers, though when asked he says it is bingo. He cannot learn new faces and do es not rememb er `pre-morbid' faces such as that of his wife and his children... Boswell can play checkers, tie his sho es, carry on a conversation, etc. Of considerable interest is the survival of Boswell's checkers skills. Evidently what we shall later term `fast' skills are 390 D. MICHIE AND R. CAMACHO Table 15.2. Of the world's three largest exp ert systems the two latest (GASOIL and BMT) were not constructed from rules obtained in dialogue fashion, but by automated induction from exp ert-supplied data. In each case the induction engineer trained the system in the desired skill in the style that the master of a craft trains an apprentice, by a structured sequence of selected examples. Rates of co de pro duction are typically in excess of 100 lines of installed Fortran, C, Pascal, etc., p er programmer day. The metho dology allows validation to b e placed on a user-transparent basis (Michie 1989), and maintenance costs are in many cases trivialized. Tabulation is from Slo comb e et al. (1986) with 1990 data on BMT added. The BMT program is describ ed on p.10 of Pragmatica, vol. 1 (ed. J.E. Hayes Michie), Glasgow, UK: Turing Institute Press. APPLICATION NO. OF DEVELOP. MAINTENANCE INDUCTIVE RULES MAN-YRS MAN-YRS/YR TOOLS medical MYCIN diagnosis 400 100 N/A N/A VAX XCON computer 8,000 180 30 N/A conguration hydro carb on separation Exp ertEase GASOIL system 2,800 1 0.1 and conguration Extran 7 conguration of 1st Class BMT re-protection > 30,000 9 2.0 and equipment in RuleMaster buildings 391 REAL-TIME SKILLS not the only ones for which pro cedural knowledge may dominate over declarative. In contrast to chess skill, checkers was already known not to lend itself to the planning approach and to b e essentially `intuitive'. When A.L. Samuel was engaged in his classic studies of machine learning using the game of checkers, he had numerous sessions with leading checkers masters directed towards dialogue acquisition of their rules and principles. Samuel rep orted (p ersonal communication) that he had never had such frustrating exp eriences in his life. In terms of relationship to what the masters actually did, the verbal material which he elicited contained almost nothing which he could use or interpret. In similar vein, Feigenbaum and McCorduck (lo c. cit., p.82) describ e this type of exp ert resp onse in the following terms: `That's true, but if you see enough patient/rocks/chipdesigns/instrument readings, you see that it is not true after all.' They conclude `At this p oint, knowledge threatens to b ecome ten thousand sp ecial cases.' The message from clinical studies is that skilled p erformance of even sophisticated tasks can still b e manifested, and learned, when the brain is so damaged that knowledge of new happ enings cannot b e retained and previously stored facts and relations (declarative-semantic memory) are seriously disrupted. Another circumstance under which the mediation of declarative memory is at least equally disabled can b e observed in the normal brain by imp osing a suciently restrictive constraint on the time available for the recognize{act cycle, as in touch-typing. This skill do es not dep end on the storage and retrieval of declarative knowledge, and can b e acquired and executed in its virtually complete absence. Recall that when copy-typing at sp eed the typist do es not need to understand the words as he or she reads them. Indeed, after a sp eed test little or nothing of the text's content can b e recalled. Moreover, educated onlo okers are surprised, although they should not b e, by the outcome of a request to the typist (supp osing that he or she has b een using a typ ewriter with unlab elled keys) to lab el the keyb oard correctly with the prop er alphanumeric symb ols. Lacking a declarative mo del, the touch-typist is ordinarily unable to do so (see, for example, Posner, 1973), other than by delib erately typing a symb ol and 392 D. MICHIE AND R. CAMACHO observing where the nger went! Simon (in press) has recently re-emphasized that simple recognition of a familiar ob ject takes at least 500 milliseconds. Op erations involving reference to a semantic mo del of the task domain require retrieval from long-term memory of relatively complex knowledge-structures and an asso ciated apparatus for inferring, storing, and utilizing intermediate results. Such elab orate transactions are to b e found only in the `slow lane'. Here seconds, minutes, or even hours are required to incubate a decision. The bare b ones of an explicit rationale for a slow-lane decision, when it comes, can usually b e elicited from the exp ert by verbal rep ort. Not so in the fast lane, to which the present discussion is conned. `Fast' skills cannot b e accessed by `dialogue elicitation' metho ds. How then are exp ert systems to b e built for these skills? A solution is to record b ehavioural traces from the exp ert sub ject. Inductive inference then reconstructs from recorded decision-data rule-based mo dels of the brain's hidden strategies. As rep orted in this review, machine execution of data-derived mo dels has b een found to generate p erformance exceeding in reliability the trained sub ject's own. 2.3. Postulates of skill acquisition Exp erimental work which will now b e describ ed was animated by a point of view ab out brains, summarized b elow as a list of p ostulates. Declarative knowledge is abbreviated to `D' and pro cedural to `P'. P designates only pro cedural knowledge which has already reached the automatized stage. 1. human agents are able verbally to rep ort their own D; 2. human agents cannot verbally rep ort their P; 3. D can b e augmented by b eing told, and also by deduction; 4. P is built by learning, whether by imitation or by trial and error; 5. P can b e executed indep endent of D, but not vice versa; 6. decision-taking via P is fast relative to use of D; 7. suciently fast control skills dep end on P alone; 393 REAL-TIME SKILLS 0500 1000 1500 2000 0 -0.01 0.01 0.02 0.03 Yaw Total Propellant used 1.66 units Yaw (degrees) Time (seconds) ωr (deg/sec) r ω 15 10 5 0 -5 Figure 15.1. Plot over time of vehicle's yaw b ehaviour (see text) controller. But it only required 1.7 units of prop ellant, a substantial saving. Also calculations and simulations by the client's engineers made the result app ear slightly b etter than optimal. This doubtless arose from minor approximations and/or distributional assumptions made in their numerical work. Time did not p ermit the p oint to b e elucidated. But the broad conclusion was seen as extremely encouraging. An industrial-strength problem had shown that the simplicity, robustness, and conceptual transparency of rule-based control does not have to b e purchased at the cost of signicant degradation of p erformance. 4. EXPERIMENTS WITH SKILL-GRAFTING Supp orted by the freedom interactively to test each conjectured mo dication on the simulator, Sammut and Michie found their own p owers of inductive conjecture adequate. But tasks of higher complexity, such as remote control of pilotless aircraft, would demand a less primitive approach. Present ideas are oriented towards the industry's use of interactive simulators for training pilots. A simulator-trained p erformer cannot tell you his or her strategy, but can demonstrate it. What is demonstrated can b e automatically recorded. What is recorded can b e inductively analysed by computer. With psychology-trained col400 D. MICHIE AND R. CAMACHO leagues, Michael Bain, Jean Hayes-Michie, and Chris Rob ertson, one of us (D.M.) engaged in an investigation into the use of the rule-induction algorithm C4.5 (see Quinlan, 1987) to uncover eective control rules from such b ehavioural records. Exp erimental sub jects were trained on an interactive simulation of a task illustrated in Figure 15.2. Control was exercised through a joystick of a p ole-and-cart simulation which refreshed the screen approximately 20 times p er second. New results together with earlier ndings with this exp erimental system (Chamb ers and Michie 1969) lead to conclusions as follows (details are available in Michie, Bain, and Hayes-Michie 1990). 4.1. Conclusions from pole-balancing First conclusion: role of problem representation. Chamb ers and Michie used two regimes of training, identical except for the graphical animation seen by the sub ject. In one variant the picture was as shown. In the other the sub ject saw only a display of four separate horizontal lines, along each of which a p ointer wandered to and fro. The sub jects in this second variant were kept in ignorance of the nature of the simulated physical system. Unknown to them, the pointers actually represented the current status of four state variables, namely p osition of cart, velocity of cart, angle of p ole, and angular velo city of p ole. Our hyp othesis was that when the system is run fast, leaving only time for use and up-dating of pro cedural memory, then there will b e no dierence in the learning curves of sub jects using the two dierent representations. Although not explicitly rep orted in their pap er, an indication of this was observed by Chamb ers and Michie. In recent work a rate was additionally used suciently slow for sub jects to rep ort the task as having a ma jor `planning' comp onent. This slow-trained group learned more slowly, at least in the initial stages. In the new work trials have not yet b een made of the lines-and-p ointers representation. Second conclusion: induction of rules from behaviour. Machine learning by imitation of a trained human was rst shown for the inverted p endulum by Donaldson (1960) and partially repro duced under bang-bang conditions by Widrow and Smith (1964). Our concern was to test the ability of mo dern induction 401 REAL-TIME SKILLS      r e - Arrow indicates current direction of motor Figure 15.2. Diagram of the pole and cart task algorithms to extract from the behavioural record the kinds of rules b elieved to accumulate in pro cedural memory during skilllearning. Results have b een p ositive. A task was investigated where the ob ject was to cross the centre of the track as often as p ossible in an allotted time-span without dropping the p ole or crashing the cart. When induction-extracted rules were installed in the computer as an `auto-pilot', p erformance on the task was similar to that of the trained human who had generated the original b ehavioural trace, but more dep endable, as describ ed b elow. Third conclusion: the clean-up eect. Rules induced from a b ehavioural record can b e assessed in two dierent ways. Predictive mo de tests the ability of a rule-set correctly to predict other b ehaviour sampled from the same source. Performance mo de tests the ability of the rules to substitute for the human source in executing the skilled task. Induced rule-sets p erformed satisfactorily in the second mo de while consistently showing high prediction error, often exceeding 20 per cent. One of the team, Mr. Michael Bain, p ointed out that when watching a machine-generated rule-set's p erformance on the screen one is struck by an app earance of sup erhuman precision and stability. A trained human skill, although controlled by an equally precise and stable set of pro duction rules, is obliged to execute via an error-prone sensorimotor system. Inconsistency and moments of inattention would then b e stripp ed away by the averaging eect implicit in inductive generalization, thus restoring to the exp erimenters a cleaned-up 402 D. MICHIE AND R. CAMACHO Table 15.4. Clean-up eect shown by induced control rules over a 5-minute test p erio d. x = p osition,  = angle: `dot' denotes rst derivatives. These results are typical, and have b een many times conrmed in test runs with the same, and with other, sub jects. x _x  _  Trained human (ranges) 2.79 4.85 0.562 5.021 Induced rule (ranges) 0.46 1.83 0.134 2.276 Range dierences 2.33 3.02 0.428 2.745 `Clean-up' 83% 62% 76% 55% version of the original pro duction rules. When tested in predictive mode, such a rule-set can do no b etter than the cumulative sum of human perceptual and execution errors allow. But in p erformance mo de one would exp ect a sup er-reliable stereotyp e of the b ehaviour of the human exemplar. Direct conrmation of this idea was obtained by calculating the magnitude of the pole and cart's excursions during a control session along each of the four dimensions of the state space. Observed ranges tabulated in Table 15.4 were obtained from a b ehavioural trace recorded from Mr. Bain's own trained performance. The ndings suggest that `skill-grafting' from behavioural traces may b e p ossible for more demanding tasks, such as those encountered in aircraft ight control. The key idea is that if we could lo ok inside the head of the ground-based pilot of a remotely controlled aircraft, or of the on-b oard pilot of a dicult vehicle such as a helicopter, we might see a neural enco ding of a fully sucient skill, but degraded in real-time execution by sensorimotor delays and errors. Recovery of a logically equiva403 REAL-TIME SKILLS lent rule structure and its transplantation to an error-free device (i.e. to a control computer) then oers a source of enhanced and more reliable p erformance. In advanced rotorcraft control there is a current need for libraries of individual autopilot manouevres (`circle at 50 feet', `y slowly sideways for one minute', etc.) which the pilot could activate in dicult weather or other conditions, so as to free his attention for some main task in hand, visual search of water surface, target acquisition, etc. 4.2. Learning to fly Sammut and colleagues have recently b een able to repro duce the `skill-grafting' phenomenon in the complex task of ying a simulated aircraft (Sammut, Hurst, Kedzier, and Michie, 1992). Using a ight simulator develop ed by Silicon Graphics, three sub jects trained themselves by rep eatedly piloting a simulated Cessna through the successive stages of a dened ight plan, consisting of the following manouevres: 1. Take o and y to an altitude of 2000 feet. 2. Level out and y to a distance of 32 000 feet from the starting p oint. 3. Turn right to a compass heading of approximately 330  . 4. At a North/South distance of 42 000 feet, turn left to head back towards the runway. 5. Line up on the runway. 6. Descend to the runway, keeping in line. 7. Land on the runway. Taking `events' as b eing signalled by the o ccurrence of control actions, then up to 1000 events were recorded p er ight. Each of three trained sub jects p erformed 30 ights, so that the complete data comprised ab out 90 000 events. For each event the control action was recorded, together with values of state variables measured at a moment selected 1{3 seconds earlier. The `oset' makes approximate allowance for the pilot's delay in resp onding to complex stimuli. To give a rough impression of the data, the following are names of recorded variables: b o olean variables: on-ground, g-limit, wing-stal l ; 404 D. MICHIE AND R. CAMACHO twist, elevation, azimuth, rol l-speed, integer variables: climb-speed, fuel, thrust, aps ; real variables: E/W distance, altitude, N/S distance, rol lers, elevator . The simulation program was mo died to log the sub jects' actions during ight. Log les from trained sub jects were used to create the input to an inductive rule-learning program. The learning program was Quinlan's (1987) C4.5. Its output to ok the form of separate decision trees for each of the four dierent control actions, further sub-divided into the seven stages listed ab ove. For example, to quote from the original pap er, The critical rule at take-o is the elevator rule: elevation > 4: level-pitch elevation  4 airsp eed  0: level-pitch airsp eed > 0: pitch-up-5 This states that as thrust is applied and the elevation is level, pull back on the stick until the elevation increases to 4  . Because of the delay, the nal elevation usually reaches 11  which is close to the values usually obtained by the pilot. `pitch-up-5' indicates a large elevator action, whereas `pitchup-1' would indicate a gentle elevator action. The other signicant control at this stage is aps: elevation  6: full-aps elevation > 6: no-aps Once the aircraft has reached an elevation angle of 6  , the aps are raised. The 28 decision trees were automatically converted to C-co de routines, arranged as a suite of seven ight control mo dules, each resp onsible for all asp ects of a given stage. A new mo dule was 405 REAL-TIME SKILLS invoked as so on as a pre-programmed precondition was satised for the onset of the next stage. Within each module, four sets of if{then rules separately sup ervized the four separate control actions. An autopilot was generated in this fashion from each of the trained sub jects. Tests were made by running the simulator in autopilot mo de, substituting as autopilot co de one or another of the three inductively synthesized program suites. The entire ight plan was executed with conspicuous comp etence, but with individual mannerisms characteristic of the ying styles of the individual human data source. Indications of the `clean-up eect' (see earlier) were also evident, particularly during the approach stage. 4.3. Learning to fly straight What is the signicance of the foregoing exp eriment? Primarily that a suitable decomp osition of the problem allows the skillgrafting metho dology to b e scaled up. Inductive skill-grafting evidently is not just applicable to p ole-balancing but also to more complex domains such as ight control. The same workers also rep orted indications of the `clean-up' eect earlier found in the p ole-balancing exp eriments, but these indications were of a preliminary nature only. We now rep ort a more detailed examination of this phenomenon indep endently conducted by Camacho (1992) using a more challenging ight control task. He used a computer simulation (ACM publicdomain software down-loaded onto a Sun Sparcstation 2) of the F-16 combat aircraft. Using Quinlan's C4.5 (see Quinlan, 1987) decision-tree induction package Camacho not only found that clean-up was op erating, but was also able to show that in his exp erimental context it played a very large, almost dominating, role. Camacho followed a similar metho dology to that of Sammut et al. , details b eing as follows. Flight plan stages: 1. Take o. 406 D. MICHIE AND R. CAMACHO 2. Climb to 1500 feet. 3. Reduce climbing angle and thrust attaining level ight at 2 kilo-feet. 4. Fly parallel to the runway's long axis for a distance of 200 kilo-feet. 5. Turn left 270  . 6. Turn right to line up with the runway. 7. As so on as distance to runway is less than 70 kilo-feet, start descent to runway keeping in line. 8. Land on the runway. Variables sampled. real: magnitude of airsp eed (knots) (Geoparallel system) real: y co ordinate of airsp eed (knots) (Geoparallel system) integer: x p osition (ft) (Geoparallel system) integer: y p osition (ft) (Geoparallel system) integer: altitude (ft) real: climb rate (ft/h) real: g-force vector in acft system (only z co ordinate) real: roll rate (rad/sec) real: pitch rate (rad/sec) real: yaw rate (rad/sec) real: heading (rad) /* Euler angles for acft */ real: pitch (rad) /* Euler angles for acft */ real: roll (rad) /* Euler angles for acft */ real: angle of attack (rad) real: angle of sideslip (rad) real: elevators setting (radl) real: ailerons setting (rad) real: rudder setting (rad) real: elevator trim setting (NOT used) real: aps setting (rad) real: sp eedBrake setting (rad) integer: throttle b o olean: gear handle b o olean: brakes 407 REAL-TIME SKILLS Figure 15.3. The `clean-up' eect. Plotted lines show distances travelled in the horizontal plane from take-o by human pilot (light lines) and the autopilot (heavy line) using the ACM ight simulator of the F-16 combat plane (see text): the y axis represents deviations in the horizontal plane from straight ight. b o olean: afterBurner Control commands used were: elevators, rollers, rudder, aps, sp eed brake angle, throttle, gear handle (b o olean), brakes (b o olean), after burner (b o olean). Thus for each of the ight plan's eight stages nine separate decision trees were synthesized. From each of typically twenty missions, successive `statevectors' were sampled and written to le, making ab out 213 000 `state-records' in all. As a p ost-pro cessing op eration, b etween 1100 and 1600 `events' were then machine-selected from each of these, making ab out 25 000 `events-records'. As in Sammut et al. , only those state vectors were selected which precede by a xed interval in the le the subsequent record of a control action. The set of events so constructed formed the `training set' for inductive synthesis of a complete autopilot of the form: ight plan plus 72 decision trees. The earlier-mentioned clean-up eect b ecame evident when 408 D. MICHIE AND R. CAMACHO Figure 15.4. Further measurements of the `clean-up' eect, see previous gure. The six thin-line plots represent the rst six missions of a total of 20 own by the human pilot (R. Camacho) to form the training set of ab out 25 000 events. The general app earance of the remaining 14 was very similar when plotted in the same way (see text): x represents distance travelled in the horizontal plane from take-o; the y axis represents deviations in the horizontal plane from straight ight. 409 REAL-TIME SKILLS matching a class of functions , in Pro c. Third Internat. Conf. on Medic. Electronics, 173-178. Feigenbaum, E.A. and McCorduck, P. (1983). The Fifth Generation: Articial Intelligence and Japan's Computer Challenge to the World, Reading, MA: Addison-Wesley. French, R.M. (1990). Subcognition and the limits of the Turing Test . In Mind, 99 , pp.53-65. Ginsb erg, M.L. (ed. 1987). Readings in Nonmonotonic Reasoning , Los Altos, CA: Morgan Kaufman. McCarthy, J. (1959). Programs with common sense . In Mechanization of Thought Pro cesses Vol. I , London: Her Ma jesty's Stationary Oce. Reprinted in M. Minsky, ed. (1960), Semantic Information Processing , Cambridge, MA: MIT Press. Makarovic, A. (1987). Pole-balancing as a benchmark problem for qualitative model ling . Technical Rep ort DP-4953, Ljubljana: Josef Stefan Institute. Revised as (1991): A qualitative way of solving the p ole-balancing problem. In Machine Intelligence, 12 (eds. J.E. Hayes, D. Michie and E. Tyugu), Oxford University Press. Michie, D. (1988). The Fifth Generation's unbridged gap . In A HalfCentury of the Universal Turing Machine (ed. R. Herken), Oxford University Press. Michie, D. (1989). Problems of computer-aided concept formation . In Applications of Exp ert Systems 2 (ed. R. J. Quinlan). Wokingham and Reading, MA: Addison Wesley, pp 310-333. Michie, D. (1991). Methodologies from machine learning in data analysis and software . Computer Journal, 34 , 559-565. Michie, D., Bain, M. and Hayes-Michie, J. E. (1990). Cognitive models from subcognitive skil ls . In Knowledge-based Systems in Industrial Control (eds. Grimble,M., McGhee, S. and Mowforth, P.), Peter Peregrinus. Michie, D. and Chambers, R.A. (1968) BOXES: an experiment in adaptive control , In Machine Intelligence 2 (eds. E. Dale and D. Michie), Edinburgh: Edinburgh University Press. Newell, A. and Simon, H.A. (1976). Computer science as empirical inquiry: symbols and search . Commun. of the ACM, 19 , 113-126. Nii, H. P. (1986). Blackboard systems: the blackboard model of problem solving and the evolution of blackboard architectures . In AI Magazine, 7 , 38-53. 416 D. MICHIE AND R. CAMACHO Popp er, K.R. and Eccles, J.C. (1977). The Self and its Brain . London and New York: Routledge and Kegan Paul. Posner, M.I. (1973). Cognition: An Introduction , Glenview, IL: Scott, Foresman. Quinlan, J.R. (1987). Generating production rules from decision trees . Internat. Joint Conf. on Art. Intell. 1987 (IJCAI-87), Los Altos, CA: Kaufmann, pp.304-307. Sammut, C. (1988). Experimental results from an evaluation of algorithms that learn to control dynamic systems . In Pro c. Fifth Internat. Conf. on Machine Learning (ed. J. Laird), San Mateo: Morgan Kaufmann. Sammut, C. (1992). Automatical ly constructing control systems by observing human behaviour . In Pro c. of the Internat. Workshop on Inductive Logic Programming, June 6-7, 1992, Tokyo, Japan. Sammut, C., Hurst, S., Kedzier, D. and Michie, D. (1992). Learning to y . In Pro c. Ninth Intern. Machine Learning Conf. (eds. D.H. Sleeman and P. Edwards), San Mateo, CA: Morgan Kaufman, pp. 385-393. Sammut, C. and Michie, D. (1991). Control ling a `black box' simulation of a space craft . The AI Magazine, Amer. Asso c. for Artif. Intell., Vol. 12 (Part 1, Spring), pp. 56-63. Simon, H.A (in press) Machine as mind , In Pro ceedings of the Turing 1990 Collo quium, 3-6 April 1990, Brighton, UK (ed. Millican, P.) to app ear. Slo comb e, S., Mo ore, K. and Zelouf, M. (1986). Engineering expert system applications . Presented at BCS Annual Conference, Decemb er 1986. Squire, L. R. (1987). Memory and Brain , Oxford University Press. Sterling, L. and Shapiro, E. (1986). The Art of Prolog , Cambridge, MA: The MIT Press, p.357. Widrow, B. and Smith, F.W (1964). Pattern recognising control systems , In Computer and Information Sciences (eds. Tou, J.T. and Wilcox, R.H.). Clever Hume Press. 417