Buffer Overflow in Programming

Lecture



Buffer overflow is a phenomenon that occurs when a computer program writes data beyond the boundaries of a buffer allocated in memory.

A buffer overflow usually arises from improper handling of externally supplied data and of memory, in the absence of strict protection from the programming subsystem (the compiler or interpreter) and the operating system. As a result of an overflow, the data located after the buffer (or before it) may be corrupted .

Buffer overflow is one of the most popular ways of breaking into computer systems , because most high-level languages use the stack frame technique — placing data on the process stack, mixing program data with control data (including the address of the start of the stack frame and the return address of the function being executed).

A buffer overflow can cause a program to crash or hang, leading to a denial of service (DoS). Certain kinds of overflows, for example an overflow in a stack frame, allow an attacker to load and execute arbitrary machine code on behalf of the program and with the privileges of the account under which it is running .

There are known examples in which a buffer overflow is deliberately used by system programs to get around limitations in existing software or hardware-software tools. For example, the iS-DOS operating system (for ZX Spectrum computers) exploited a buffer overflow in the built-in TR-DOS to launch its own loader in machine code (something that cannot be done in TR-DOS by standard means).

Security

A program that uses a vulnerability to break the protection of another program is called an exploit. The greatest danger comes from exploits designed to gain superuser-level access or, in other words, to escalate privileges. A buffer overflow exploit achieves this by passing the program specially crafted input data. Such data overflows the allocated buffer and modifies the data that follows that buffer in memory.

Consider a hypothetical system administration program that runs with superuser privileges — for example, one that changes user passwords. If the program does not check the length of the new password entered, any data whose length exceeds the size of the buffer allocated to store it will simply be written over whatever came after the buffer. An attacker can insert into this memory area machine-language instructions, for example shellcode, that perform any actions with superuser privileges — adding and deleting user accounts, changing passwords, modifying or deleting files, and so on. If execution is permitted in this memory area and the program later transfers control to it, the system will execute the attacker's machine code located there.

Properly written programs should check the length of the input data to make sure it is no larger than the allocated data buffer. However, programmers often forget to do this. If the buffer is located on the stack and the stack "grows downward" (for example, on the x86 architecture), then a buffer overflow can be used to change the return address of the function being executed, since the return address is located after the buffer allocated by the executing function. This makes it possible to execute an arbitrary piece of machine code in the address space of the process. It is possible to use a buffer overflow to corrupt the return address even if the stack "grows upward" (in which case the return address is usually located before the buffer).

Even experienced programmers can find it difficult to determine how far a particular buffer overflow may be a vulnerability. This requires deep knowledge of the computer architecture and of the target program. It has been shown that even overflows as small as writing a single byte beyond the boundary of the buffer can constitute vulnerabilities.

Buffer overflows are widespread in programs written in relatively low-level programming languages such as assembly language, C and C++, which require the programmer to manage the size of allocated memory manually. Eliminating buffer overflow errors is still a poorly automated process. Formal program verification systems are not very effective with modern programming languages.

Many programming languages, such as Perl, Python, Java and Ada, manage memory allocation automatically, which makes errors related to buffer overflow unlikely or impossible. To avoid buffer overflows, Perl automatically resizes arrays. However, the runtime systems and libraries for such languages may still be susceptible to buffer overflows because of possible internal errors in the implementation of these checking systems. On Windows, some software and hardware-software solutions are available that prevent the execution of code beyond an overflowed buffer if such an overflow has occurred. Among these solutions are DEP in Windows XP SP2, OSsurance and Anti-Execute.

In the Harvard architecture, executable code is stored separately from data, which makes such attacks practically impossible.

Brief technical overview

Example

Consider an example of a vulnerable program in C:

  Buffer Overflow in Programming

It uses the unsafe function strcpy, which allows more data to be written than the array allocated for it can hold. If this program is run on a Windows system with an argument longer than 100 bytes, the program will most likely terminate abnormally, and the user will receive an error message.

The following program is not subject to this vulnerability:

  Buffer Overflow in Programming

Here strcpy is replaced by strncpy, in which the maximum number of characters copied is limited by the size of the buffer.[11]

Description

The diagrams below show how a vulnerable program can damage the structure of the stack.

Illustration of writing various data to a buffer allocated on the stack

  • Buffer Overflow in Programming

    A. — Before the data is copied.

  • Buffer Overflow in Programming

    B. — The string "hello" has been written to the buffer.

  • Buffer Overflow in Programming

    C. — The buffer is overflowed, which has overwritten the return address.

On the x86 architecture, the stack grows from higher addresses to lower ones, that is, new data is placed before the data already on the stack.

By writing data to the buffer, one can write beyond its boundaries and modify the data located there, in particular, change the return address.

If the program has special privileges (for example, it is run with root rights), an attacker can replace the return address with the address of shellcode, which will allow the attacker to execute commands on the attacked system with elevated privileges.[12]

Exploitation

Techniques for exploiting a buffer overflow vary depending on the architecture, the operating system and the memory region. For example, a buffer overflow in the heap (used for dynamic memory allocation) differs significantly from the analogous case on the call stack.

Stack exploitation

Also known as stack smashing. A technically skilled user can use a stack buffer overflow to take control of a program for their own purposes in the following ways:

  • by overwriting a local variable located in memory next to the buffer, changing the behavior of the program to the attacker's advantage.
  • by overwriting the return address in the stack frame. As soon as the function finishes, control is transferred to the address specified by the attacker, usually into a memory region that the attacker had access to modify.
  • by overwriting a function pointer or an exception handler, which will subsequently receive control.
  • by overwriting a parameter from another stack frame or a non-local address that is referenced in the current context.

If the address of the user data is unknown but is stored in a register, the "trampolining" technique can be used: the return address can be overwritten with the address of an opcode that will transfer control to the memory region containing the user data. If the address is stored in register R, then jumping to an instruction that transfers control to that address (for example, call R) will cause the user-specified code to execute. The addresses of suitable opcodes or memory bytes can be found in DLLs or in the executable file itself. However, the addresses usually cannot contain null characters, and the locations of these opcodes vary depending on the application and the operating system. The Metasploit Project, for example, maintained a database of suitable opcodes for Windows systems (it is currently unavailable).

A stack buffer overflow should not be confused with a stack overflow.

It is also worth noting that such vulnerabilities are usually found using the testing technique known as fuzzing.

Heap exploitation

A buffer overflow in the heap data area is called a heap overflow and is exploited differently from a stack buffer overflow. Heap memory is allocated dynamically by the application at run time and usually contains program data. Exploitation is carried out by corrupting this data in particular ways so as to make the application overwrite internal structures, such as pointers in linked lists. The usual exploit technique for a heap buffer overflow is to overwrite dynamic memory links (for example, malloc metadata) and use the resulting modified pointer to overwrite a function pointer of the program.

A vulnerability in Microsoft's GDI+ product that arises when processing JPEG images is an example of the danger that a heap buffer overflow can pose.[16]

Difficulties in exploitation

Manipulating the buffer before it is read or executed can hinder the successful exploitation of a vulnerability. Such manipulations can reduce the threat of a successful attack, but not eliminate it entirely. They may include converting a string to upper or lower case, removing special characters, or filtering out everything except alphanumeric characters. However, there are techniques for getting around these measures: alphanumeric shellcodes,[17] polymorphic[18] and self-modifying code, and the return-to-library attack. The same methods can be used to hide from intrusion detection systems. In some cases, including those involving conversion of characters to Unicode, a vulnerability is mistakenly assumed to allow only a DoS attack, when in fact remote execution of arbitrary code is possible.

Prevention

Various techniques are used to make a buffer overflow less likely.

Intrusion detection systems

Intrusion detection systems (IDS) can detect and prevent attempts at remote exploitation of a buffer overflow. Since in most cases the data intended to overflow a buffer contains long sequences of No Operation instructions (NOP or NOOP), an IDS simply blocks all incoming packets containing a large number of consecutive NOPs. This method is, in general, ineffective, since such sequences can be written using a variety of assembly language instructions. Recently, crackers have begun to use shellcode with encryption, self-modifying code, polymorphic code and alphanumeric code, as well as return-to-standard-library attacks, in order to get past IDS.

Stack corruption protection

Stack corruption protection is used to detect the most common buffer overflow errors. It checks that the call stack has not been modified before returning from a function. If it has been modified, the program terminates with a segmentation fault.

There are two systems, StackGuard and the Stack-Smashing Protector (formerly called ProPolice), both of which are extensions of the gcc compiler. Beginning with gcc-4.1-stage2, SSP was integrated into the main distribution of the compiler. Gentoo Linux and OpenBSD include SSP in the gcc distributed with them

Placing the return address on the data stack makes it easier to carry out a buffer overflow that leads to the execution of arbitrary code. In theory, gcc could be modified so that the address is placed in a special return stack completely separate from the data stack, similar to how it is implemented in the Forth language. However, this is not a complete solution to the buffer overflow problem, since other stack data also needs protection.

Executable space protection for UNIX-like systems

Executable space protection can mitigate the consequences of buffer overflows by making most of an attacker's actions impossible. It is achieved through address space layout randomization (ASLR) and/or by forbidding simultaneous write and execute access to memory. A non-executable stack prevents most shellcode exploits.

There are two patches for the Linux kernel that provide this protection — PaX and exec-shield. Neither of them has yet been included in the mainline kernel. Starting with version 3.3, OpenBSD includes a system called W^X, which also provides control over executable space.

Note that this method of protection does not prevent stack corruption. However, it often prevents the successful execution of the exploit's "payload". The program will not be able to insert shellcode into write-protected memory, such as existing executable code segments. It will also be impossible to execute instructions in non-executable memory, such as the stack or the heap.

ASLR makes it harder for an attacker to determine the addresses of functions in the program code that could be used to carry out a successful attack, and makes ret2libc-type attacks a very difficult task, although they are still possible in a controlled environment or if the attacker guesses the required address correctly.

Some processors, such as Sun's Sparc, Transmeta's Efficeon, and the newest 64-bit processors from AMD and Intel, prevent the execution of code located in memory regions marked with a special NX bit. AMD calls its solution NX (from No eXecute), and Intel calls its own XD (from eXecute Disabled).

Executable space protection for Windows

There are currently several different solutions designed to protect executable code on Windows systems, offered both by Microsoft and by third-party companies.

Microsoft proposed its own solution, called DEP (Data Execution Prevention), and included it in service packs for Windows XP and Windows Server 2003. DEP uses additional features of new Intel and AMD processors that were designed to overcome the 4 GB limit on addressable memory inherent in 32-bit processors. For these purposes, some service structures were enlarged. These structures now contain a reserved NX bit. DEP uses this bit to prevent attacks involving modification of the exception handler address (the so-called SEH exploit). DEP protects only against the SEH exploit; it does not protect memory pages containing executable code.

In addition, Microsoft developed a stack protection mechanism intended for Windows Server. The stack is marked with so-called "canaries", whose integrity is then checked. If a canary has been modified, the stack has been corrupted.

There are also third-party solutions that prevent the execution of code located in memory regions intended for data, or that implement an ASLR mechanism.

Using safe libraries

The buffer overflow problem is characteristic of the C and C++ programming languages, because they do not hide the details of the low-level representation of buffers as containers for data types. Therefore, to avoid buffer overflows, a high level of control must be maintained over the creation and modification of the program code that manages buffers. Using libraries of abstract data types that perform centralized automatic buffer management and include overflow checking is one engineering approach to preventing buffer overflow.

The two main data types that allow a buffer overflow to occur in these languages are strings and arrays. Thus, using libraries for strings and list data structures that were designed to prevent and/or detect buffer overflows makes it possible to avoid many vulnerabilities. The price of such solutions is reduced performance due to the extra checks and other actions performed by the library code, since it is written "for all occasions", and in any particular case some of the actions it performs may be redundant.

History

Buffer overflow was understood and partially documented as early as 1972 in the publication "Computer Security Technology Planning Study". The earliest documented malicious use of a buffer overflow occurred in 1988. One of the several exploits used by the Morris worm to propagate itself across the Internet was based on it. The program exploited a vulnerability in the finger service of the Unix system. Later, in 1995, Thomas Lopatic independently rediscovered the buffer overflow and posted the results of his research to the Bugtraq list A year later, Elias Levy published a step-by-step introduction to exploiting stack buffer overflows, "Smashing the Stack for Fun and Profit", in Phrack magazine.

Since then, at least two well-known network worms have used buffer overflows to infect a large number of systems. In 2001, the Code Red worm exploited this vulnerability in Microsoft's Internet Information Services (IIS) 5.0, and in 2003 SQL Slammer infected machines running Microsoft SQL Server 2000.

In 2003, the use of a buffer overflow present in licensed Xbox games made it possible to run unlicensed software on the console without hardware modification, using so-called modchips. The PS2 Independence Exploit also used a buffer overflow to achieve the same result for the PlayStation 2. A similar exploit for the Wii, Twilight Hack, used this vulnerability in the game The Legend of Zelda: Twilight Princess.

See also

  • Heap overflow
  • Stack overflow
  • Off-by-one error
  • Heap spraying

Comments

To leave a comment

If you have any suggestion, idea, thanks or comment, feel free to write. We really value feedback and are glad to hear your opinion.
To reply

Lectures and tutorial on "Algorithmization and programming. Structural programming. C language"

Terms: Algorithmization and programming. Structural programming. C language