Lecture
Friendly artificial intelligence (also friendly AI or FAI) is a hypothetical artificial general intelligence (AGI) that would have a positive (beneficial) effect on humanity, or at least align with human interests or contribute to fostering the improvement of the human species. It is part of the ethics of artificial intelligence and is closely related to machine ethics. While machine ethics is concerned with how an artificially intelligent agent should behave, friendly artificial intelligence research focuses on how to practically bring about such behavior and ensure it is adequately constrained.
This term was coined by Eliezer Yudkowsky , who is best known for popularizing the idea of discussing superintelligent artificial agents that reliably implement human values. In the leading artificial intelligence textbook by Stuart J. Russell and Peter Norvig, «Artificial Intelligence: A Modern Approach», the idea is described:
Yudkowsky (2008) elaborates further on how to design a friendly AI. He argues that friendliness (a desire not to harm humans) should be designed in from the start, but that designers must recognize that their own designs may be flawed, and that the robot will learn and evolve over time. Thus the challenge is one of mechanism design—to define a mechanism for evolving AI systems under a system of checks and balances, and to give the systems utility functions that will remain friendly in the face of such changes.
«Friendly» is used in this context as technical terminology, and picks out agents that are safe and useful, not necessarily ones that are «friendly» in the colloquial sense. The concept is invoked primarily in the context of discussions of recursively self-improving artificial agents that rapidly explode in intelligence, on the grounds that this hypothetical technology would have a large, rapid, and difficult-to-control impact on human society.
The roots of concern about artificial intelligence are very old. Kevin LaGrandeur has shown that the dangers specific to AI can be seen in ancient literature concerning artificial humanoid servants such as the golem, or the proto-robots of Gerbert of Aurillac and Roger Bacon. In these stories, the exceptional intelligence and power of these humanlike creations clash with their status as slaves (which are by nature regarded as sub-human), and cause catastrophic conflict. By 1942 these themes had prompted Isaac Asimov to create the «Three Laws of Robotics»—principles hard-wired into every robot in his fiction, meant to prevent them from turning on their creators or allowing them to come to harm.
In modern times, as the prospect of superintelligent AI draws ever closer, philosopher Nick Bostrom has said that superintelligent AI systems with goals that are not aligned with human ethics are intrinsically dangerous unless extreme measures are taken to ensure the safety of humanity. He put it this way:
Basically we should assume that a «superintelligence» would be able to achieve whatever goals it has. Therefore, it is extremely important that the goals we endow it with, and its entire motivation system, be «human friendly».
In 2008, Eliezer Yudkowsky called for the creation of «friendly AI» to reduce the existential risk posed by advanced artificial intelligence. He explains: «The AI does not hate you, nor does it love you, but you are made out of atoms which it can use for something else».
Steve Omohundro says that a sufficiently advanced AI system, if not explicitly counteracted, will exhibit a number of basic «drives», such as resource acquisition, self-preservation, and continuous self-improvement, owing to the intrinsic nature of any goal-driven systems, and that these drives will, «without special precautions», cause the AI to behave in undesirable ways.
Alexander Wissner-Gross says that AIs that seek to maximize their future freedom of action (or causal path entropy) can be considered friendly if their planning horizon is longer than a certain threshold, and unfriendly if their planning horizon is shorter than that threshold.
Luke Muehlhauser, writing for the Machine Intelligence Research Institute, recommends that machine ethics researchers adopt what Bruce Schneier has called the «security mindset»: rather than thinking about how a system will work, imagine how it might fail. For example, he suggests that even an AI that only makes accurate predictions and communicates through a text interface could cause unintended harm.
In 2014, Luke Muehlhauser and Nick Bostrom underscored the need for «friendly AI»; nonetheless, the difficulties in designing a «friendly» superintelligence, for instance by programming counterfactual moral reasoning, are considerable.
Yudkowsky advances the model of coherent extrapolated volition (CEV). In his words, our coherent extrapolated volition is «our wish if we knew more, thought faster, were more the people we wished we were, had grown up farther together; where the extrapolation converges rather than diverges, where our wishes cohere rather than interfere». ; extrapolated as we wish to extrapolate, interpreted as we wish to interpret».
Rather than being designed directly by human programmers, Friendly AI is meant to be designed by a «seed AI» programmed first to study human nature, and then to produce the AI that humanity would want, given enough time and insight, to arrive at a satisfactory answer.[16] Appealing to a goal via contingent human nature (perhaps expressed, for mathematical purposes, in the form of a utility function or other decision-theoretic formalism) as providing the ultimate criterion of «Friendliness» is a response to the meta-ethical problem of defining an objective morality; extrapolated volition is intended to be what humanity would objectively want, all things considered, but it can only be defined relative to the psychological and cognitive qualities of present-day, unextrapolated humanity.
Steve Omohundro has proposed a «scaffolding» approach to AI safety, in which one provably safe generation of AI helps build the next provably safe generation.
Seth Baum argues that the development of safe, socially beneficial artificial intelligence or artificial general intelligence is a function of the social psychology of the communities engaged in AI research, and so can be constrained by extrinsic measures and motivated by intrinsic measures. Intrinsic motivation can be strengthened when messages resonate with AI developers; Baum argues that, by contrast, «existing messages about beneficial AI are not always well framed». Baum advocates for «cooperative relationships and a positive framing of AI researchers», and cautions against characterizing AI researchers as «unwilling to pursue beneficial designs».
In his book «Human Compatible », AI researcher Stuart J. Russell lists three principles that should guide the design of beneficial machines. He stresses that these principles are not meant to be explicitly coded into the machines; rather, they are intended for human designers. The principles are as follows:
1. The machine's only objective is to maximize the realization of human preferences.
2. The machine is initially uncertain about what those preferences are.
3. The ultimate source of information about human preferences is human behavior.
The «preferences» Russell refers to are «all-encompassing; they cover everything you might care about, arbitrarily far into the future». Likewise, «behavior» includes any choice between options, and the uncertainty is such that some probability, which may be quite small, must be assigned to every logically possible human preference.
James Barrat, author of the book «Our Final Invention», suggested that «a public-private partnership needs to be created to bring AI developers together to share ideas about security—something like the International Atomic Energy Agency, but in partnership with corporations». He calls on AI researchers to convene a meeting similar to the Asilomar Conference on Recombinant DNA, at which the risks of biotechnology were discussed.
John McGinnis urges governments to accelerate friendly AI research. Since the goals of friendly AI need not be outstanding, he proposes a model similar to that of the National Institutes of Health, where «panels of peers in computer and cognitive science would sift through projects and select those that are designed both to advance AI and to ensure that such advances are accompanied by appropriate safeguards». McGinnis believes that peer review is better «than regulation for solving technical problems that cannot be solved through bureaucratic authority». McGinnis notes that his proposal stands in contrast to that of the Machine Intelligence Research Institute, whose aim, generally speaking, is to avoid government involvement in friendly AI.
According to Gary Marcus, the annual amount of money spent on developing machine morality is negligible.
Some critics believe that both human-level AI and superintelligence are unlikely, and that friendly AI is therefore unlikely. Writing for The Guardian, Alan Winfield compares human-level artificial intelligence to faster-than-light travel in terms of difficulty, and states that, while we must be «cautious and prepared» given the stakes, we «needn't be obsessed» with the risks of superintelligence. Boyles and Joaquin, on the other hand, argue that Luke Muehlhauser and Nick Bostrom's proposal for creating friendly AIs seems bleak. This is because Muehlhauser and Bostrom appear to hold the idea that intelligent machines can be programmed to reason counterfactually about the moral values that people ought to have had. In an article in AI & Society, Boyles and Joaquin argue that such AIs would not be as friendly as claimed, given the following: the infinite number of prior counterfactual conditions that would have to be programmed into the machine, the difficulty of cashing out a set of moral values—that is, ones that are more ideal than those humans currently possess—and, as a consequence, the apparent mismatch between the counterfactual prior values and the ideal values.
Some philosophers claim that any truly «rational» agent, whether artificial or human, will naturally be benevolent; from this point of view, deliberate safety measures intended to create a friendly AI could be unnecessary or even harmful. Other critics wonder whether artificial intelligence can be friendly at all. Adam Keiper and Ari N. Schulman, editors of the technology journal The New Atlantis, say that it is impossible ever to guarantee «friendly» behavior in AI, because problems of ethical complexity will not yield to progress in software or to increases in computing power. They write that the criteria on which friendly AI theories are based work «only when someone possesses not only a great capacity to predict the likelihood of a host of possible outcomes, but also certainty and consensus about how the different outcomes are valued».
Comments