An Evaluation of the Remote Viewing Program: Research and Operational Applications (AIR Draft Report)
This CIA reading-room document is the American Institutes for Research (AIR) draft report, dated September 22, 1995, evaluating the intelligence community's remote viewing program prior to its potential transfer to a new sponsoring organization. Commissioned by the CIA's Office of Research and Development in June 1995 after CIA declassified its past parapsychology efforts, the review examined two components: the laboratory research program (chiefly work at Stanford Research Institute/SRI International and Science Applications International Corporation/SAIC under Principal Investigator Dr. Edwin May) and the operational use of remote viewing for intelligence collection under the DIA's 'Star Gate' program. A blue-ribbon panel including statistician Dr. Jessica Utts and skeptic Dr. Raymond Hyman, plus AIR scientists Dr. Michael Mumford and Dr. Andrew Rose, Dr. Lincoln Moses, and AIR president Dr. David Goslin, reviewed roughly 80 publications and interviewed end-users, viewers, and the program manager. The report concludes that while a statistically significant laboratory effect was observed, it could not be unambiguously attributed to a paranormal cause, and that remote viewing never produced actionable intelligence—recommending against continuation. Dr. Utts's included assessment argues, by contrast, that psychic functioning has been scientifically well established. The document was approved for release by the CIA on 2002/05/22.
Description
1995 draft evaluation report prepared for the CIA by the American Institutes for Research (AIR) assessing the U.S. intelligence community's remote viewing program (Star Gate), covering both its laboratory research and operational intelligence-gathering applications. Includes the reviews of statistician Dr. Jessica Utts and psychologist Dr. Raymond Hyman.
Claims
CIA's program began in 1972 and was discontinued in 1977; DIA's direct involvement began about 1985.
85%Dr. Jessica Utts concluded that psychic functioning has been well established and results are far beyond chance.
85%Dr. Raymond Hyman represented a skeptical position questioning the paranormal interpretation.
85%Remote viewing never produced actionable intelligence and was not warranted for continued intelligence use (AIR conclusion).
80%CIA's program began in 1972 and was discontinued in 1977; DIA involvement began about 1985.
80%It is unclear whether observed effects can be unambiguously attributed to paranormal ability rather than methodological characteristics.
75%Remote viewing never produced actionable intelligence and its continued use in intelligence gathering is not warranted (AIR conclusion).
75%It remains unclear whether observed laboratory effects can be attributed to paranormal ability rather than methodological characteristics.
75%A statistically significant laboratory effect was demonstrated, in the sense that hits occur more often than chance.
70%A statistically significant laboratory effect was demonstrated, with hits occurring more often than chance.
70%Dr. Jessica Utts concludes that psychic functioning has been well established using standards applied to any other area of science.
60%
Events
Dec 31, 1971
CIA remote viewing program begins
CIA's parapsychology program began in 1972 and was discontinued in 1977.
Jun 30, 1995
AIR interviews conducted
Interviews of end-users, viewers, and the program manager conducted during July and August 1995.
Sep 21, 1995
AIR draft report completed
American Institutes for Research completed the draft evaluation report.
Dec 31, 1987
NRC report on human performance published
National Academy Press published 'Enhancing Human Performance,' reviewing techniques including remote viewing.
Dec 31, 1989
SAIC assumes program sponsorship
Government research sponsorship moved from SRI to SAIC under Dr. Edwin May.
Dec 31, 1994
Program suspended
Operational remote viewing operations continued until Spring 1995, when the program was suspended.
Dec 31, 1994
CIA declassifies past parapsychology program
In 1995 the CIA declassified past efforts to facilitate external review; ORD contracted AIR in June 1995.
May 21, 2002
Document approved for release
CIA approved the document for release.
Dec 31, 1976
CIA program discontinued
CIA discontinued its remote viewing program.
Dec 31, 1987
NRC review published
National Academy Press published 'Enhancing Human Performance,' reviewing remote viewing largely negatively.
Dec 31, 1994
Operational program suspended
Star Gate operational remote viewing was suspended in Spring 1995.
May 21, 2002
CIA declassification release
Document approved for release by CIA.
Dates mentioned
Keywords
Entities
Organizations
Programs
Phenomena
Extracted text (OCR)
Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 An Evaluation of the Remote Vewing Program: Research and Operational Applications | * DRAFT REPORT Prepared by: The American Institutes for Research September 22, 1995 Note: The page sequence in this document has been corrected from the original record. Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Executive Summary Studies of paranormal phenomena have nearly always been associated with controversy. Despite the controversy concerning their nature and existence, many individuals and organizations continue to be avidly interested in these phenomena. The intelligence community is no exception: beginning in the 1970s, it has conducted a program intended to investigate the application of one paranormal phenomenon — remote viewing, or the ability to describe locations ong has not | visited. hee ne prin knew le hye? Conceptually, remote viewing would seem to have tremendous potential utility for the intelligence community. Accordingly, a three-component program involving basic research, operations, and foreign assessment has been in place for some time. Prior to transferring this program to a new sponsoring organization within the intelligence community, a thorough program review was initiated. The part of the program review conducted by the American Institutes for Research (AIR), a nonprofit, private research organization, consisted of two main components. The first component was a review of the research program. The second component was a review of the operational application of the remote viewing phenomenon in intelligence gathering. Evaluation of the foreign assessment component ofthe program was not within the scope of the present effort. Research Evaluation To evaluate the research program, a “blue ribbon" panel was assembled. The panel included two noted experts in the area of parapsychology: Dr. Jessica Utts, a Professor of Statistics at the University of California/Davis, and Dr. Raymond Hyman, a Professor of Psychology at the University of Oregon. In addition to their extensive credentials, they were selected to represent both sides of the paranormal controversy: Dr. Utts has published articles that view paranormal interpretations positively, while Dr. Hyman was selected to represent a more skeptical position. Both, however, are viewed as fair and open-minded scientists, In addition to these experts, this panel included two Senior Scientists from AIR; both have recognized methodological expertise, and both had no prior background in parapsychological research. They were included in the review panel to provide an unbiased methodological DRAFT - American Institutes for Research E-F Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Executive Summary perspective. In addition, Dr. Lincoln Moses, an Emeritus Professor at Stanford University, provided statistical advice, while Dr. David A, Goslin, President of AIR, served as coordinator of the research effort. Panel members were asked to review all laboratory experiments and meta-analytic reviews conducted as part of the research program; this consisted of approximately 80 Separate publications, many of which are summiary reports of multiple experiments. In the course of this review, special attention was given to those studies that (a) provided the strongest evidence for the remote viewing phenomenon, and ( b) represented new experiments controlling for methodological artifacts identified in earlier reviews. Separate written reviews were prepared by Dr. Utts and Dr. Hyman. They exchanged reviews with other panel members who then tried to reach a consensus. In the typical remote viewing experiment in the laboratory, a remote viewer is asked to visualize a place, location, or object being viewed by a "beacon" or sender. A judge then examines the viewer's report and determines if this report matches the target or, alternatively, a set of decoys. In most recent laboratory experiments reviewed for the present evaluation, National Geographic photographs provided the target pool. If the viewer's reports match the target, as opposed to the decoys, a hit is said to have occurred. Alternatively, accuracy of a set of remote viewing reports is assessed by rank-ordering the similarity of each remote viewing report to each photograph in the target set (usually five photographs). A better-than- chance score is presumed to represent the occurrence of the paranormal phenomenon of remote viewing, sirfce the remote viewers had not seen the photographs they had described (or did not know which photographs had been randomly selected for a particular remote viewing trial), In evaluating the various laboratory studies conducted to date, the reviewers reached the following conclusions: * A Statistically significant laboratory effort has been demonstrated in the sense that hits occur more often than chance, * It is unclear whether the observed effects can unambiguously be attributed to the paranormal ability of the remote viewers as opposed to characteristics of the DRAFT - American Institutes for Research E-2 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Executive Summary judges or of the target or some other characteristic of the methods used. Use of the same remote viewers, the same judge, and the same target photographs makes it impossible to identify their independent effects. + Evidence has not been provided that clearly demonstrates that the causes of hits are due to the operation of paranormal phenomena; the laboratory experiments have not ~~ identified the sources or origins of the remote viewing phenomenon. Operational Evaluation The second component of the program involved the use of remote viewing in gathering intelligence information. Here, representatives of various intelligence groups— “end users" of intelligence information — presented targets to remote viewers, who were asked to describe the target. Typically, the remote viewers described the results of their experiences in written reports, which were forwarded to the end users for evaluation and, if warranted, action. To assess the operational value of remote viewing in intelligence gathering, a multifaceted evaluation strategy was employed. First, the relevant research literature was reviewed to identify whether the conditions applying during intelligence gathering would _ reasonably permit application of the remote viewing paradigm. Second, members of three groups involved in the program were interviewed: (1) end users of the information; (2) the remote viewers providing the reports, and (3) the program manager. Third, feedback information obtained from end user judgments of the accuracy and value of the remote viewing reports were assessed. This multifaceted evaluation effort led to the following conclusions: ¢ The conditions under which the remote viewing phenomenon is observed in laboratory settings do not apply in intelligence gathering situations. For example, viewers cannot be provided with feedback and targets may not display the v characteristics needed to produce hits. DRAFT - American Institutes for Rasearch E-3 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Executive Summary ¢ The end users indicating that, although some accuracy was observed with regard to broad background characteristics, the remote viewing reports failed to produce the concrete, specific information valued in intelligence gathering. ¢ The information provided was inconsistent, inaccurate with regard to specifics, and required substantial subjective interpretation. ° In no case had the information provided ever been uSed to guide intelligence operations. Thus, Remote viewing failed to produce actionable intelligence. Conclusions ce The foregoing observations provide a compelling argument against continuation of the J an program within the intelligence community. Even though a statistically significant effect has been observed in the laboratory, it remains unclear whether the existence of a paranormal phenomenon, remote viewing, has been demonstrated. The laboratory studies do not provide t-— evidence regarding the sources or origins of the phenomenon, nor do they address an important methodological issue of inter-judge reliability. Further, even if it could be demonstrated unequivocally that a paranormal phenomenon occurs under the conditions present in the laboratory paradigm, these conditions have limited applicability and utility for intelligence gathering operations. For example, the nature of the remote viewing targets are vastly dissimilar, as are the specific tasks required of a the remote viewers. Most importantly, the information provided by remote viewing is vague and ambiguous, making it difficult, if not impossible, for the technique to yield information of sufficient quality and accuracy fof iforzatigl for actionable intelligence. Thus, we conclude that continued use of remote viewing in intelligence gathering operations is not warranted. DRAFT - American institutes for Research E-g Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 1. Background and History In their continuing quest to improve effectiveness, many organizations have sought techniques that might be used to enhance performance. For the most part, the candidate techniques cme from rather traditional lines of inquiry stressing interventions such as selection, training, and performance appraisal. However, some other, more controversial performance enhancement techniques have also been suggested. These techniques range from implicit learning and mental rehearsal to the enhancement of paranormal abilities. In the mid-1980s, at the request of the Army Research Institute, the National Research Council of the National Academy of Sciences established a blue ribbon panel charged with evaluating the evidence bearing on the effectiveness of a wide variety of techniques for enhancing human performance. This review was conducted under the overall direction of David A. Goslin, then Executive Director of the Commission on Behavioral and Social Sciences and Education (CBASSE), and now President of the American Institutes for Research (AIR). The review panel's report, Enhancing Human Performance: Issues, Theories, and Techniques, was published by the National Academy Press in 1988 and summarized by Swets and Bjork (1990). They noted that although the panel found some support for certain alternative performance enhancement techniques — for example, guided imagery — little or no support was found for the usefulness of many other techniques, such as learning during sleep and remote viewing. Although the findings of the National Research Council (NRC) were predominantly negative with regard to a range of paranormal phenomena, work on remote viewing has continued under the auspices of various government programs. Since 1986, perhaps 50 to 100 additional studies of remote viewing have been conducted. At least some of these studies represent significant attempts to address the methodological problems noted in the review conducted by the NRC panel. At the request of Congress, the Central Intelligence Agency (CIA) is considering assuming responsibility for this the remote viewing program. As part of its decision-making process, the CIA was asked to evaluate the research conducted since the NRC report. This evaluation was intended to determine: (a) whether this research has any long-term practical The American Institutes for Research 7-7 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter One: Background and History value for the intelligence community, and (b) if it does, what changes should be made in methods and approach to enhance the value of remote viewing research. To achieve these goals, the CIA contracted with the American Institutes for Research to supervise and conduct the evaluation. This report contains the results of our evaluation. Before presenting our results, we begin by presenting a brief overview of the remote viewing phenomenon and a short history of the applied program that involves remote viewing. Remote Viewing Although parapsychological research has a long history, studies of "remote viewing" — also referred to as a form of "anomalous cognition" — as a unique manifestation of psychic functioning began in the 1970s. In its simplest form, a typical remote viewing study during this early period of investigation consisted of the following: A p: rson, referred to as a "beacon" or "sender," travels to a series of remote sites. The remote viewer, a person who putatively has the parapsychological ability, is asked to describe the locations of the beacon. Typically, these location descriptions include drawings and a verbal description of the location. Subsequently, a judge evaluates this description by rank ordering the set of locations against the descriptions. If the judge finds that the viewer's description most closely matched the actual location of the sender, a hit is said to have occurred. If hits occur more often than chance, or if the assigned ranks are more accurate than a random assignment, one might argue that a psychic phenomenon has been observed: the viewer has described a location not visited during the session. This phenomenon has been studied by various investigators throughout the intervening period, using several variants of this basic paradigm. If certain people (or all people to a greater or lesser extent, as has been proposed by some investigators) possess the ability to see and describe target locations they have not visited, this ability might prove of great value to the intelligence community. As an adjunct method to gathering intelligence, people who possess this ability could be asked to describe various intelligence targets. This information, especially if considered credible and reliable, could supplement and enhance more time-consuming and perhaps dangerous methods for collecting data. Although certain (perhaps unwarranted) assumptions, such as the availability of a sender, are implicit in this argument, the possibility of gathering intelligence through this mechanism has provided the major impetus for government interest in remote viewing. The American Institutes for Research 1-2 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter One: Background and History Remote viewing was and continues to be a controversial phenomenon. Early research on remote viewing was plagued by a number of statistical and methodological flaws.’ One statistical flaw found in early studies of remote viewing, for example, was due to failure to control for the elimination of locations already judged. For example, if there were five targets in the set, judges might lower their rankings for a viewing already judged as a "hit" or ranked first. In other words, all targets did not have an equal probability of being assigned all ranks. Another commonly noted methodological flaw was that cues in the remote viewing paradigm, such as the time needed to drive to various locations, may have allowed viewers to produce hits without using any parapsychological ability. More recent research has attempted to control for many of these problems. New paradigms have been developed where, for example, viewers — in double-blind conditions — are asked to visualize pictures drawn from a target pool consisting of National Geographic photographs. In addition to this experimental work, an applied program of intelligence operations actually using remote viewers has been developed. In the following section, we describe the history of the government's remote viewing program. Program History "Star Gate" is a Defense Intelligence Agency (DIA) program which involved the use of paranormal phenomena, primarily “remote viewing," for intelligence collection. During Star Gate's history, DIA pursued three basic program objectives: "Operations," using remote viewing to collect intelligence against foreign targets; "Research and Development," using laboratory studies to find new ways to improve remote viewing for use in the intelligence world; and "Foreign Assessment," the analysis of foreign activities to develop or exploit the paranormal for any uses which might affect our national security. Prior to the advent of Star Gate in the early 1990s, the DIA, the Central Intelligence Agency (CIA), and other government organizations conducted various other programs pursuing some or all of these objectives, CIA's program began in 1972, but was discontinued in 1977, DIA's direct involvement began about 1985 and has continued up to the time of this 1Many of these problems are described in the National Research Council Report. DRAFT - American Institutes for Research 1- Some, Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter One: Background and History review. During the last twenty years, all government programs involving parapsychology have been viewed as highly controversial, high-risk, and have been subjected to various reviews. In 1995, the CIA declassified its past parapsychology program efforts in order to facilitate a new, external review. In addition, CIA worked with DIA to continue declassification of Star Gate program documents, a process which had already begun at DIA. All relevant CIA an DIA program documents were collected and inventoried. In June of 1995, CIA's Office of Research and Development (ORD) then contracted with AIR for this external review, based on our long-standing expertise in carrying out studies relating to behavioral science issues and our neutrality with respect to the subject matter. Evaluation Objectives The CIA asked AIR to address a number of key objectives during the technical review of Star Gate. These included: ° a comprehensive evaluation of the research and development in this area, with a focus on the validity of the technical approach(es) according to acceptable scientific standards. ° an evaluation of the overall program utility or usefulness to the government. The CIA believes that the controversial nature of past parapsychology programs within the intelligence community, and the scientific controversy clouding general acceptance of the validity of paranormal phenomena, demand that these two issues of utility and scientific validity be addressed separately. . consideration of whether any changes in the operational or research and development activities of the program might bring about improved results if the results were not already optimum. The American Institutes for Research 1-4 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter One: Background and History ° develop recommendations for the CIA as to appropriate strategies for program activity in the future. We were directed to base our findings on the data and information provided as a result of DIA and CIA program efforts, since it was neither possible nor intended that we review the entire field of parapsychological research and its applications. Also, we would not review or evaluate the "Foreign Assessment" component of the program. In the next chapter, we present our methodology for conducting the evaluation. A major component of the evaluation was to commission two nationally-regarded experts to review the program's relevant research studies; their findings are presented in Chapter 3, along with our analysis of areas of agreement and disagreement. In Chapter 4, we present our findings concerning the operational component of the program. Finally, in Chapter 5 we present our conclusions and recommendations. The American institutes for Research 1-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 et Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 2. Evaluation Plan The broad goal of the present effort was to provide a thorough and objective evaluation of the remote viewing program. Because of the multiple components of the program, a multifaceted evaluation plan was devised. As mentioned previously, only the research and intelligence gathering components of the program were considered here. In this section, we describe the general approach used in evaluating these two components of the program, beginning with the research program. Remote Viewing Research The Research Program. The government-sponsored research program had three broad objectives. The first and primary objective was to provide scientifically compelling evidence for the existence of the remote viewing phenomenon. It could be argued that if unambiguous evidence for the existence of the phenomenon cannot be provided, then there is little reason to be concerned with its potential applications. The second objective of the research program was to identify causal mechanisms that might account for or explain the observed (or inferred) phenomenon. This objective of the program is of some importance; an understanding of the origins of a phenomenon provides a basis for developing potential applications, Further, it provides more compelling evidence for the existence of the phenomenon (Cook & Campbell, 1979; James, Muliak, & Brett, 1982). Thus, in conducting a thorough review, an attempt must be made to assess the success of the program in developing an adequate explanation of the phenomenon. The third objective of the research program was to identify techniques or procedures that might enhance the utility of the information provided by remote viewings. For example, how might more specific information be obtained from viewers and what conditions set boundaries on the accuracy of viewings? Research along those lines is of interest primarily because it provides the background necessary for operational applications of the phenomenon, The NRC provided a thorough review of the unclassified remote viewing research through 1986. In this review (summarized in Swets & Bjork, 1990), the nature of the research methods led the reviewers to question whether there was indeed any effect that could DRAFT - American Institutes for Research 2-1 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter Two: Evaluation Plan clearly be attributed to the operation of paranormal phenomena. Since then, the Principal Investigator, Dr, Edwin May, under formerly classified government contracts, has conducted a number of other studies not previously reviewed. These studies were expressly intended to address many of the criticisms raised in the initial NRC report. Because these studies might provide new evidence for the existence of the remote viewing phenomenon, its causal mechanisms, and its boundary conditions, a new review seemed called for. The Review Panel. With these issues in mind, a blue-ribbon review panel was commissioned, with the intent of ensuring a balanced and objective appraisal of the research. Two of the reviewers were scientists noted for their interest, expertise, and experience in parapsychological research. The first of these two expert reviewers, Dr. Jessica Utts, a Professor of Statistics at the University of California-Davis, is a nationally-recognized scholar who has made major contributions to the development and application of new statistical methods and techniques. Among many other positions and awards, Dr. Utts is an Associate Editor of the Journal of the American Statistical Association (Theory and Methods) and the Statistical Editor of the Journal of the American Society for Psychical Research. She has published several articles on the application of statistical methods to parapsychological research and has direct experience with the remote viewing research program. The second expert reviewer, Dr. Raymond Hyman, is a Professor of Psychology at the University of Oregon, Dr. Hyman has published over 200 articles in professional journals on perception, pattern recognition, creativity, problem solving, and critiques of the paranormal, He served on the original NRC Committee on Techniques for the Enhancement of Human Performance. Dr. Hyman serves as a resource to the media on topics related to the paranormal, and has testified as an expert witness in court cases involving paranormal claims. He is recognized as one of the most important and fair-minded skeptics working in this area, Curriculum Vitae for Dr. Utts and Dr. Hyman are included in Appendix A. In addition to these two experts, four other scientists were involved in the work of the review panel. Two senior behavioral scientists and experts in research methods at the American Institutes for Research, Dr. Michael Mumford and Dr. Andrew Rose, served both as members of and staff to the panel. Dr. Mumford holds a Ph.D. in Industrial/Organizational Psychology from the University of Georgia. He is a Fellow of the American Psychological Association's Division 5, Measurement, Evaluation, and Statistics. Dr. Rose is a cognitive DRAFT - American institutes for Research 2-2 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter Two: Evaluation Plan psychologist with a Ph.D. from the University of Michigan. He has over 22 years of experience in designing and conducting basic and applied behavic. al science research. Dr. Rose is Chief Scientist of the Washington Office of AIR. They were to bring to the panel a methodological perspective unbiased by prior work in the area of parapsychology. The third participant was Dr, Lincoln Moses, an Emeritus Professor of Statistics at Stanford University, who participated in the review as a resource with regard to various statistical issues. Finally, Dr. David A. Goslin, President of ATR, participated as both a reviewer and coordinator for the review panel. Research Content. Prior to convening the first meeting of the review panel, the CIA transferred to AIR all reports and documents relevant to the review. We organized and copied these documents. In addition, the Principal Investigator for the program, Dr. Edwin May, was asked to provide two other pieces of information for the panel. First, he was asked to list those studies which he believes provide the strongest evidence bearing on the nature and significance of the remote viewing phenomenon. Second, he was asked to identify all unique studies conducted since the initial NRC report that provide evidence bearing on the nature and significance of the phenomenon. Additionally, he was asked to participate in an interview with members of the review panel following its first meeting to clarify any ambiguities about these studies. The complete list of documents, including notations of the "strongest evidence" set and the "unique" set, is included in Appendix B.' Review Procedures. Remote viewing, like virtually all other parapsychological phenomena, represents one of the most controversial research areas in the social sciences (e.g., Bem & Honorton, 1994; Hyman, 1994). Therefore, any adequate review of the research program must take this controversy into account in such a way that the review procedures are likely to result in a fair and unbiased assessment of the research. To ensure a fair and comprehensive review, Drs, Utts and Hyman agreed to examine all program documents. In the course of this review it was agreed that all members of the review panel would carefully consider: 1One document pertaining to the program remained classified during the period of this review. One of the review panel (Dr. Mumford) examined this document and provided an unclassified synopsis to the review panel. DRAFT - American institutes for Research 2-3 Approved For Release 2002/05/22 : CIA-RDP96-00791 R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter Two: Evaluation Plan ° those studies recommended by the Principal Investigator as providing compelling evidence for the phenomenon, and ° those empirical studies conducted since the NRC review that might provide new evidence about the existence and nature of the phenomenon. The members of the review panel convened at the Palo Alto office of AIR to structure exactly how the review process would be carried out. To ensure that different perspectives on paranormal phenomena would be adequately represented, Drs. Utts and Hyman were asked to prepare independent reports based on their review. In this review, they were to cover four general topics: ° Was there a statistically significant effect? ° Could the observed effect, if any, be attributed to a paranormal phenomenon? ° What mechanisms, if any, might plausibly be used to account for any significant effects and what boundary conditions influence these effects? . What would the findings obtained in these studies indicate about the characteristics and potential applications of information obtained through the remote viewing process? After they had each completed their reports, they presented the reports to other members of the panel. After studying these reports, all members of the review panel (except Dr. Moses) participated in a series of conference calls, The primary purpose of these exchanges was to identify the conclusions on which the experts agreed and disagreed. Next, in areas where they disagreed, Drs. Utts and Hyman were asked to discuss the nature of the disagreements, determine why they disagreed, and if possible, attempt to resolve the disagreements. Both the initial reports and the dialogue associated with discussion of any disagreements were made a part of the written record. In fact, Dr. Hyman's opinions on areas of agreement and disagreement are included in his report; in addition to her initial report, Dr. Utts prepared a reply to Dr. Hyman's opinions of agreement and disagreement. This reply, in addition to their original reports, are included in Section 3 below. DRAFT - American institutes for Research 2-4 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter Two: Evaluation Plan If disagreements could not be resolved through this dialogue, then the other members of the review panel were to consider the remaining issues from a general methodological perspective. Subsequently, they were to provide an addendum to the dialogue indicating which of the two positions being presented seemed to be on firmer ground both substantively and methodologically. This addendum concludes Section 3 below. Intelligence Gathering: The Operational Program The Program. In addition to the research component, the program included two operational components. One of those components was "foreign assessment," or analysis of the paranormal research being conducted by other countries. This issue, however, is beyond the scope of the present review. The other component involved the use of remote viewing as a technique for gathering intelligence information. In the early 1970s, the CIA experimented with applications of remote viewing in intelligence gathering. Later in the decade, they abandoned the program. However, other government agencies, including the Department of Defense, used remote viewers to obtain intelligence information. The viewers were tasked with providing answers to questions posed by various intelligence agencies. These operations continued until the Spring of 1995, when the program was suspended. Although procedures varied somewhat during the history of the program, viewers typically were presented with a request for information about a target of interest to a particular agency. Multiple viewings were then obtained for the target. The results of the viewings then were summarized in a three- or four-page report and sent to the agency that had posed the original question. Starting in 1994, members of the agencies receiving the viewing reports were formally asked to evaluate their accuracy and value. Any comprehensive evaluation of the remote viewing program must consider how viewings were used by the intelligence community. One might demonstrate the existence of a statistically significant paranormal phenomenon in experiments conducted in the laboratory; however, the phenomenon could prove to be of limited operational value either because it does not occur consistently outside the laboratory setting or because the kind of information DRAFT « American Institutes for Research 2-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter Two: Evaluation Plan provided is of limited value to the intelligence community. General Evaluation Procedures. No one piece of evidence provides unequivocal support for the usefulness of a program. Instead, a more accurate and comprehensive picture can be obtained by considering multiple sources of evidence (Messick, 1989). Three basic sources of information were used in evaluation of the intelligence gathering component:. ° Prior research studies . Interviews with program participants, and . Analyses of user feedback. Prior Research Studies. As noted above, one aspect of the laboratory research program was to identify those conditions that set bounds on the accuracy and success of the remote viewing process. Thus, one way to analytically evaluate potential applications in intelligence gathering is to enumerate the conditions under which viewers were assigned tasks and then examine the characteristics of the remote viewing paradigm as studied through experimentation in the laboratory. The conditions under which operational tasks occur — that is, the requirements imposed by intelligence gathering — could then provide an assessment of the applicability of the remote viewing process. Interviews. As part of the Star Gate program, the services of remote viewers were used to support operational activities in the intelligence community. This operational history provides an additional basis for evaluating the Star Gate program; ultimately, if the program is to be of any real value, it must be capable of serving the needs of the intelligence community. By examining how the remote viewing services have been used, it becomes possible to draw some initial, tentative conclusions about the potential value of the Star Gate program. Below, we describe how information bearing on intelligence applications of the remote viewing phenomenon was gathered. Later, in Section 4, we describe the results of this information-gathering activity and draw some conclusions from the information we obtained. Although a variety of techniques might be used to accrue retrospective information (questionnaires, interviews, diaries, etc.), the project team decided that structured interviews DRAFT - American Institutes for Research 2-6 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter Two: Evaluation Plan examining issues relevant to the various participants would provide the most appropriate strategy. Accordingly, structured interviews were developed for three participant groups in intelligence operations: ° end-users: representatives from agencies requesting information from remote viewers, ° the Program Manager, and ° the remote viewers. Another key issue to be considered in an interview procedure is the nature of the people to be interviewed. Although end-users, program managers, and viewers represent the major participants, many different individuals have been involved in intelligence applications of remote viewing over the course of the last twenty years. Nevertheless, it was decided to interview only those persons who were involved in the program at the time of its suspension in the Spring of 1995. This decision was based on the need for accurate, current information that had not been distorted by time and could be corroborated by existing documentation and follow-up interviews. Information about operational applications was gathered in a series of interviews conducted during July and August of 1995. We interviewed seven representatives of end-user groups, three remote viewers, and the incumbent Program Manager. With regard to the data collection procedures that we employed, a number of points should be borne in mind. First, members of the groups we interviewed could only speak to recent operations. Although it would have been desirable to interview people involved in earlier operations, for example during the 1970s, the problems associated with the passage of time, including forgetting and the difficulties involved in verifying information, effectively precluded this approach. Accordingly, the interviews focused on current operations. Second, it should be noted that the end-user representatives represented a range of current concerns in the intelligence community. The relevant user groups were involved in operations ranging from counterintelligence and drug interdiction to search and rescue operations. This diversity permitted operational merits to be assessed for a number of different contexts. DRAFT - American institutes for Research 2-/ Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter Two: Evaluation Plan The interviews were conducted by one of the two panel members from AIR. A retired intelligence officer took notes during the interviews. A representative of the CIA attended interviews as necessary to describe the reasons the interviews were being conducted and to address any security concerns. Each interview was conducted using a standard protocol. Different protocols were developed for members of the three groups because they had somewhat different perspectives on current operations. Appendix C presents the instructions given to the interviewer. This Appendix also lists the interview questions presented to users, viewers, and the program manager. User interviews were conducted in the offices of the client organization; interviews with the program manager and the viewers were conducted at the Washington Office of AIR. The interviews were one to two hours long. A total of 12 to 16 questions were asked in the interviews. We developed the questions presented in each interview,as follows: Initially, the literature on remote viewing and available information bearing on operations within the intelligence community were reviewed by AIR scientists. This review was used to formulate an initial set of interview questions. Subsequently, these candidate questions were presented to a panel of three psychologists at AIR. In addition, review panel members were asked to review these candidate questions to insure they were not leading and covered the issues that were relevant to the particular group under consideration. _ With regard to operational users, four types of questions were asked. These four types of questions examined the background and nature of the tasks presented to the remote viewers, the nature and accuracy of the information resulting from the viewings, operational use of this information, and the utility of the resulting information. The remote viewers were asked a somewhat different set of questions. The four types of questions presented to them examined recruitment, selection, and development; the procedures used to generate viewings; the conditions that influenced the nature and success of viewings; and the organizational factors that influenced program operations. The Program Manager was not asked about the viewing process. Instead, questions presented to the program manager primarily focused on broader organizational issues. The DRAFT - American Institutes for Research 2-8 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter Two: Evaluation Plan four types of managerial questions focused on the manager's background, client recruitment, factors influencing successes and failures, and needs for effective program management. The interview questions presented in each protocol were asked in order, as specified in Appendix D. Typically these interviews began by asking for objective background information. Questions examining broader evaluative issues were asked at the end of the interview. The AIR scientist conducting the interviews produced reports for each individual interview. They also are contained in Appendix C. Analyses of User Feedback. In addition to the qualitative data provided by the interviews, some quantitative information was available. For all of the operational tasks conducted during 1994, representatives from the requesting agencies were asked to provide two summary judgments: one with respect to the accuracy of the remote viewing, and the second of the actual or potential value of the information provided. These data — the accuracy and value evaluations obtained for viewings as program feedback from the users — were analyzed and summarized in a report prepared prior to the current evaluation. A copy of this report is provided in Appendix D. Although these judgments have been routinely collected for only a relatively short period of time, they provided an important additional source of evaluative information. This information was of some value as a supplement to interviews in part because it was collected prior to the start of the current review, and in part because it reflects user assessments of the resulting information. We present the findings flowing from this multifaceted evaluation of the operational component of the program in Section 4 of this report. In that section, we first present the findings emerging from prior research and the interviews and then consider the results obtained from the more quantitative evaluations. Prior to turning to this evaluation of operations, however, we first present the findings from review of the basic research, examining evidence for the existence and nature of the remote viewing phenomenon. DRAFT - American Institutes for Research 2-9 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 3. Research Reviews In this section, we present the conclusions drawn by the two experts after reviewing the research studies bearing on remote viewing. We begin by presenting the review of Dr. Jessica Utts. Subsequently, a rejoinder is provided by Dr. Raymond Hyman. Finally, Dr. Utts presents a reply to Dr. Hyman. The major points of agreement and disagreement are noted in the final section, along with our conclusions. In conducting their reviews, both Dr. Hyman and Dr. Utts focused on the remote viewing research. However, additional material is provided as indicated by the need to clarify "certain points being made. Furthermore, both reviewers provided unusually comprehensive reviews considering not only classified program research, but also a number of earlier studies having direct bearing on the nature and significance of the phenomenon. An Assessment of the Evidence for Psychic Functioning Dr. Jessica Utts Division of Statistics, University of California, Davis September 1, 1995 ABSTRACT Research on psychic functioning, conducted over a two decade period, is examined to determine whether or not the phenomenon has been scientifically established. A secondary question is whether or not it is useful for government purposes. The primary work examined in this report was government sponsored research conducted at Stanford Research Institute, later known as SRI International, and at Science Applications International Corporation, known as SAIC. Using the standards applied to any other area of science, it is concluded that psychic functioning has been well established. The statistical results of the studies examined are far beyond what is expected by chance. Arguments that these results could be due to methodological flaws in the experiments are soundly refuted. Effects of similar magnitude to those found in government-sponsored research at SRI and SAIC have been replicated at a number of laboratories across the world. Such consistency cannot be readily explained by claims of flaws or fraud. DRAFT - American Institutes for Research 3-1 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter Three: Research Reviews The magnitude of psychic functioning exhibited appears to be in the range between what social scientists call a small and medium effect. That means that it is reliable enough to be replicated in properly conducted experiments, with sufficient trials to achieve the long-run statistical results needed for replicability. A number of other patterns have been found, suggestive of how to conduct more productive experiments and applied psychic functioning. For instance, it doesn't appear that a sender is needed. Precognition, in which the answer is known to no one until a future time, appears to work quite well. Recent experiments suggest that if there is a psychic sense then it works much like our other five senses, by detecting change. Given that physicists are currently grappling with an understanding of time, it may be that a psychic sense exists that scans the future for major change, much as our eyes scan the environment for visual change or out ears allow us to respond to sudden changes in sound. It is recommended that future experiments focus on understanding how this phenomenon works, and on how to make it as useful as possible. There is little benefit to continuing experiments designed to offer proof, since there is little more to be offered to anyone who does not accept the current collection of data. 1. INTRODUCTION The purpose of this report is to examine a body of evidence collected over the past few decades in an attempt to determine whether or not psychic functioning is possible. Secondary questions include whether or not such functioning can be used productively for government purposes, and whether or not the research to date provides any explanation for how it works. There is no reason to treat this area differently from any other area of science that relies on statistical methods. Any discussion based on belief should be limited to questions that are not data-driven, such as whether or not there are any methodological problems that could substantially alter the results. It is too often the case that people on both sides of the question debate the existence of psychic functioning on the basis of their personal belief systems rather than on an examination of the scientific data. American Institutes for Research 3-2 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter Three: Research Reviews One objective of this report is to provide a brief overview of recent data as well as the scientific tools necessary for a careful reader to reach his or her own conclusions based on that data. The tools consist of a rudimentary overview of how statistical evidence is typically evaluated, and a listing of methodological concerns particular to experiments of this type. Government-sponsored research in psychic functioning dates back to the early 1970s when a program was initiated at what was then the Stanford Research Institute, now called SRI International. That program was in existence until 1989. The following year, government sponsorship moved to a program at Science Applications International Corporation (SAIC) under the direction of Dr. Edwin May, who had been employed in the SRI program since the mid 1970s and had been Project Director from 1986 until the close of the program. This report will focus most closely on the most recent work, done by SAIC. Section 2 describes the basic statistical and methodological issues required to understand this work; Section 3 discusses the program at SRI; Section 4 covers the SAIC work (with some of the details in an Appendix); Section 5 is concerned with external validation by exploring related results from other laboratories; Section 6 includes a discussion of the usefulness of this capability for government purposes and Section 7 provides conclusions and recommendations. 2. SCIENCE NOTES 2.1 DEFINITIONS AND RESEARCH PROCEDURES There are two basic types of functioning that are generally considered under the broad heading of psychic or paranormal abilities. These are classically known as extrasensory perception (ESP), in which one acquires information through unexplainable means and psychokinesis, in which one physically manipulates the environment through unknown means. The SAIC laboratory uses more neutral terminology for these abilities; they refer to ESP as anomalous cognition (AC) and to psychokinesis as anomalous perturbation (AP). The vast majority of work at both SRI and SAIC investigated anomalous cognition rather than anomalous perturbation, although there was some work done on the latter. American Institutes for Research 3-3 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter Three: Research Reviews Anomalous cognition is further divided into categories based on the apparent source of the information. If it appears to come from another person, the ability is called telepathy, if it appears to come in real time but not from another person it is called clairvoyance and if the information could have only been obtained by knowledge of the future, it is called precognition. It is possible to identify apparent precognition by asking someone to describe something for which the correct answer isn't known until later in time. It is more difficult to rule out precognition in experiments attempting to test telepathy or clairvoyance, since it is almost impossible to be sure that subjects in such experiments never see the correct answer at some point in the future. These distinctions are important in the quest to identify an explanation for anomalous cognition, but do not bear on the existence issue. The vast majority of anomalous cognition experiments at both SRI and SAIC used a technique known as remote viewing. In these experiments, a viewer attempts to draw or describe (or both) a target location, photograph, object or short video segment. All known channels for receiving the information are blocked. Sometimes the viewer is assisted by a monitor who asks the viewer questions; of course in such cases the monitor is blind to the answer as well. Sometimes a sender is looking at the target during the session, but sometimes there is no sender. In most cases the viewer eventually receives feedback in which he or she learns the correct answer, thus making it difficult to rule - out precognition as the explanation for positive results, whether or not there was a sender. Most anomalous cognition experiments at SRI and SAIC were of the free-response type, in which viewers were simply asked to describe the target. In contrast, a forced-choice experiment is one in which there are a small number of known choices from which the viewer must choose. The latter may be easier to evaluate statistically but they have been traditionally less successful than frec-response experiments. Some of the work done at SAIC addresses potential explanations for why that might be the case. 2.2 STATISTICAL ISSUES AND DEFINITIONS Few human capabilities are perfectly replicable on demand. For example, even the best hitters in the major baseball leagues cannot hit on demand. Nor can we predict when American Institutes for Research 3-4 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter Three: Research Reviews someone will hit or when they will score a home run. In fact, we cannot even predict whether or not a horse run will occur in a particular game. That does not mean that home runs don't exist. . Scientific evidence in the statistical realm is based on replication of the same average performance or relationship over the long run. We would not expect a fair coin to result in five heads and five tails over each set of ten tosses, but we can expect the proportion of heads and tails to settle down to about one half over a very long series of tosses. Similarly, a good baseball hitter will not hit the ball exactly the same proportion of times in each game but should be relatively consistent over the long run. The same should be true of psychic functioning. Even if there truly is an effect, it may never be replicable on demand in the short run even if we understand how it works. However, over the long run in well-controlled laboratory experiments we should see a consistent level of functioning, above that expected by chance. The anticipated level of functioning may vary based on the individual players and the conditions, just as it does in baseball, but given players of similar ability tested under similar conditions the results should be replicable over the long run. In this report we will show that replicability in that sense has been achieved. 2.2.1 P-VALUES AND COMPARISON WITH CHANCE. In any area of science, evidence based on statistics comes from comparing what actually happened to what should have happened by chance. For instance, without any special interventions about 51 percent of births in the United States result in boys. Suppose someone claimed to have a method that enabled one to increase the chances of having a baby of the desired sex. We could study their method by comparing how often births resulted in a boy when that was the intended outcome. If that percentage was higher than the chance percentage of 51 percent over the long run, then the claim would have been supported by statistical evidence. Statisticians have developed numerical methods for comparing results to what is expected by chance. Upon observing the results of an experiment, the p-value is the answer to the following question: Jf chance alone is responsible for the results, how likely would we be to observe results this strong or stronger? If the answer to that question, i.e. the p-value is very small, then most researchers are willing to rule out chance as an explanation. In fact it is commonly accepted practice to say that if the p-value is 5 percent (0.05) or less, then we can American Institutes for Research 35 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5 Chapter Three: Research Reviews rule out chance as an explanation. In such cases, the results are said to be statistically significant. Obviously the smaller the p-value, the more convincingly chance can be ruled out. Notice that when chance alone is at work, we erroneously find a statistically significant result about 5 percent of the time. For this reason and others, most reasonable scientists require replication of non-chance results before they are convinced that chance can be ruled out. 2.2.2 REPLICATION AND EFFECT SIZES: In the past few decades scientists have realized that true replication of experimental results should focus on the magnitude of the effect, or the effect size rather than on replication of the p-value, This is because the latter is heavily dependent on the size of the study. In a very large study, it will take only a small magnitude effect to convincingly rule out chance. In a very small study, it would take a huge effect to convincingly rule out chance. In our hypothetical sex-determination experiment, suppose 70 out of 100 births designed to be boys actually resulted in boys, for a rate of 70 percent instead of the 51 percent expected by chance. The experiment would have a p-value of 0.0001, quite convincingly ruling out chance. Now suppose someone attempted to replicate the experiment with only ten births and found 7 boys, ie also 70 percent. The smaller experiment would have a p-value of 0. 19, and would not be statistically significant. If we were simply to focus on that issue, the result would appear to be a failure to replicate the original result, even though it achieved exactly the same 70 percent boys! In only ten births it would require 90 percent of them to be boys before chance could be ruled out, Yet the 70 percent rate is a more exact replication of the result than the 90 percent. Therefore, while p-values should be used to assess the overall evidence for a phenomenon, they should not be used to define whether or not a replication of an experimental result was "successful." Instead, a successful replication should be one that achieves an effect that is within expected statistical variability of the original result, or that achieves an even stronger effect for explainable reasons. A number of different effect size measures are in use in the social sciences, but in this report we will focus on the one used most often in remote viewing at SRI and SAIC. Because the American Institutes for Research 3-6 Approved For Release 2002/05/22 : CIA-RDP96-00791R000200180005-5
Related files
Connected through the network
SF-2026-001036
Proposed Downgrading of Classified Documents for Cognitive Sciences Laboratory, Volume 2 of 2
SF-2026-001002
