|
|
1.1 ! root 1: This is Info file gcc.info, produced by Makeinfo-1.43 from the input ! 2: file gcc.texi. ! 3: ! 4: This file documents the use and the internals of the GNU compiler. ! 5: ! 6: Copyright (C) 1988, 1989, 1992 Free Software Foundation, Inc. ! 7: ! 8: Permission is granted to make and distribute verbatim copies of ! 9: this manual provided the copyright notice and this permission notice ! 10: are preserved on all copies. ! 11: ! 12: Permission is granted to copy and distribute modified versions of ! 13: this manual under the conditions for verbatim copying, provided also ! 14: that the section entitled "GNU General Public License" is included ! 15: exactly as in the original, and provided that the entire resulting ! 16: derived work is distributed under the terms of a permission notice ! 17: identical to this one. ! 18: ! 19: Permission is granted to copy and distribute translations of this ! 20: manual into another language, under the above conditions for modified ! 21: versions, except that the section entitled "GNU General Public ! 22: License" and this permission notice may be included in translations ! 23: approved by the Free Software Foundation instead of in the original ! 24: English. ! 25: ! 26: ! 27: File: gcc.info, Node: Non-bugs, Prev: Bug Reporting, Up: Bugs ! 28: ! 29: Certain Changes We Don't Want to Make ! 30: ===================================== ! 31: ! 32: This section lists changes that people frequently request, but which ! 33: we do not make because we think GNU CC is better without them. ! 34: ! 35: * Checking the number and type of arguments to a function which has ! 36: an old-fashioned definition and no prototype. ! 37: ! 38: Such a feature would work only occasionally--only for calls that ! 39: appear in the same file as the called function, following the ! 40: definition. The only way to check all calls reliably is to add a ! 41: prototype for the function. But adding a prototype will ! 42: eliminate the need for this feature. So the feature is not ! 43: worthwhile. ! 44: ! 45: * Warning about using an expression whose type is signed as a shift ! 46: count. ! 47: ! 48: Shift count operands are probably signed more often than unsigned. ! 49: Warning about this would cause far more annoyance than good. ! 50: ! 51: * Warning about assigning a signed value to an unsigned variable. ! 52: ! 53: Such assignments must be very common; warning about them would ! 54: cause more annoyance than good. ! 55: ! 56: * Making bitfields unsigned by default on particular machines where ! 57: "the ABI standard" says to do so. ! 58: ! 59: The ANSI C standard leaves it up to the implementation whether a ! 60: bitfield declared plain `int' is signed or not. This in effect ! 61: creates two alternative dialects of C. ! 62: ! 63: The GNU C compiler supports both dialects; you can specify the ! 64: dialect you want with the option `-fsigned-bitfields' or ! 65: `-funsigned-bitfields'. However, this leaves open the question ! 66: of which dialect to use by default. ! 67: ! 68: Currently, the preferred dialect makes plain bitfields signed, ! 69: because this is simplest. Since `int' is the same as `signed ! 70: int' in every other context, it is cleanest for them to be the ! 71: same in bitfields as well. ! 72: ! 73: Some computer manufacturers have published Application Binary ! 74: Interface standards which specify that plain bitfields should be ! 75: unsigned. It is a mistake, however, to say anything about this ! 76: issue in an ABI. This is because the handling of plain bitfields ! 77: distinguishes two dialects of C. Both dialects are meaningful on ! 78: every type of machine. Whether a particular object file was ! 79: compiled using signed bitfields or unsigned is of no concern to ! 80: functions in any other object file, even if they access the same ! 81: bitfields in the same data structures. ! 82: ! 83: A given program is written in one or the other of these two ! 84: dialects. The program stands a chance to work on most any ! 85: machine if it is compiled with the proper dialect. It is ! 86: unlikely to work at all if compiled with the wrong dialect. ! 87: ! 88: Many users appreciate the GNU C compiler because it provides an ! 89: environment that is uniform across machines. These users would be ! 90: inconvenienced if the compiler treated plain bitfields ! 91: differently on certain machines. ! 92: ! 93: Occasionally users write programs intended only for a particular ! 94: machine type. On these occasions, the users would benefit if the ! 95: GNU C compiler were to support by default the same dialect as the ! 96: other compilers on that machine. But such applications are rare. ! 97: And users writing a program to run on more than one type of ! 98: machine cannot possibly benefit from this kind of compatibility. ! 99: ! 100: This is why GNU CC does and will treat plain bitfields in the same ! 101: fashion on all types of machines (by default). ! 102: ! 103: (Of course, users strongly concerned about portability should ! 104: indicate explicitly in each bitfield whether it is signed or not.) ! 105: ! 106: * Undefining `__STDC__' when `-ansi' is not used. ! 107: ! 108: Currently, GNU CC defines `__STDC__' as long as you don't use ! 109: `-traditional'. This provides good results in practice. ! 110: ! 111: Programmers normally use conditionals on `__STDC__' to ask whether ! 112: it is safe to use certain features of ANSI C, such as function ! 113: prototypes or ANSI token concatenation. Since plain `gcc' ! 114: supports all the features of ANSI C, the correct answer to these ! 115: questions is "yes". ! 116: ! 117: Some users try to use `__STDC__' to check for the availability of ! 118: certain library facilities. This is actually incorrect usage in ! 119: an ANSI C program, because the ANSI C standard says that a ! 120: conforming freestanding implementation should define `__STDC__' ! 121: even though it does not have the library facilities. `gcc -ansi ! 122: -pedantic' is a conforming freestanding implementation, and it is ! 123: therefore required to define `__STDC__', even though it does not ! 124: come with an ANSI C library. ! 125: ! 126: Sometimes people say that defining `__STDC__' in a compiler that ! 127: does not completely conform to the ANSI C standard somehow ! 128: violates the standard. This is illogical. The standard is a ! 129: standard for compilers that are supposed to conform. It says ! 130: nothing about what any other compilers should do. Whatever the ! 131: ANSI C standard says is relevant to the design of plain `gcc' ! 132: without `-ansi' only for pragmatic reasons, not as a requirement. ! 133: ! 134: * Undefining `__STDC__' in C++. ! 135: ! 136: Programs written to compile with C++-to-C translators get the ! 137: value of `__STDC__' that goes with the C compiler that is ! 138: subsequently used. These programs must test `__STDC__' to ! 139: determine what kind of C preprocessor that compiler uses: whether ! 140: they should concatenate tokens in the ANSI C fashion or in the ! 141: traditional fashion. ! 142: ! 143: These programs work properly with GNU C++ if `__STDC__' is ! 144: defined. They would not work otherwise. ! 145: ! 146: In addition, many header files are written to provide prototypes ! 147: in ANSI C but not in traditional C. Many of these header files ! 148: can work without change in C++ provided `__STDC__' is defined. ! 149: If `__STDC__' is not defined, they will all fail, and will all ! 150: need to be changed to test explicitly for C++ as well. ! 151: ! 152: ! 153: File: gcc.info, Node: VMS, Next: Portability, Prev: Bugs, Up: Top ! 154: ! 155: Using GNU CC on VMS ! 156: ******************* ! 157: ! 158: * Menu: ! 159: ! 160: * Include Files and VMS:: Where the preprocessor looks for the include files. ! 161: * Global Declarations:: How to do globaldef, globalref and globalvalue with ! 162: GNU CC. ! 163: * VMS Misc:: Misc information. ! 164: ! 165: ! 166: File: gcc.info, Node: Include Files and VMS, Next: Global Declarations, Prev: VMS, Up: VMS ! 167: ! 168: Include Files and VMS ! 169: ===================== ! 170: ! 171: Due to the differences between the filesystems of Unix and VMS, GNU ! 172: CC attempts to translate file names in `#include' into names that VMS ! 173: will understand. The basic strategy is to prepend a prefix to the ! 174: specification of the include file, convert the whole filename to a VMS ! 175: filename, and then try to open the file. GNU CC tries various prefixes ! 176: one by one until one of them succeeds: ! 177: ! 178: 1. The first prefix is the `GNU_CC_INCLUDE:' logical name: this is ! 179: where GNU C header files are traditionally stored. If you wish ! 180: to store header files in non-standard locations, then you can ! 181: assign the logical `GNU_CC_INCLUDE' to be a search list, where ! 182: each element of the list is suitable for use with a rooted ! 183: logical. ! 184: ! 185: 2. The next prefix tried is `SYS$SYSROOT:[SYSLIB.]'. This is where ! 186: VAX-C header files are traditionally stored. ! 187: ! 188: 3. If the include file specification by itself is a valid VMS ! 189: filename, the preprocessor then uses this name with no prefix in ! 190: an attempt to open the include file. ! 191: ! 192: 4. If the file specification is not a valid VMS filename (i.e. does ! 193: not contain a device or a directory specifier, and contains a `/' ! 194: character), the preprocessor tries to convert it from Unix syntax ! 195: to VMS syntax. ! 196: ! 197: Conversion works like this: the first directory name becomes a ! 198: device, and the rest of the directories are converted into ! 199: VMS-format directory names. For example, `X11/foobar.h' is ! 200: translated to `X11:[000000]foobar.h' or `X11:foobar.h', whichever ! 201: one can be opened. This strategy allows you to assign a logical ! 202: name to point to the actual location of the header files. ! 203: ! 204: 5. If none of these strategies succeeds, the `#include' fails. ! 205: ! 206: Include directives of the form: ! 207: ! 208: #include foobar ! 209: ! 210: are a common source of incompatibility between VAX-C and GNU CC. VAX-C ! 211: treats this much like a standard `#include <foobar.h>' directive. ! 212: That is incompatible with the ANSI C behavior implemented by GNU CC: to ! 213: expand the name `foobar' as a macro. Macro expansion should ! 214: eventually yield one of the two standard formats for `#include': ! 215: ! 216: #include "FILE" ! 217: #include <FILE> ! 218: ! 219: If you have this problem, the best solution is to modify the source ! 220: to convert the `#include' directives to one of the two standard forms. ! 221: That will work with either compiler. If you want a quick and dirty ! 222: fix, define the file names as macros with the proper expansion, like ! 223: this: ! 224: ! 225: #define stdio <stdio.h> ! 226: ! 227: This will work, as long as the name doesn't conflict with anything else ! 228: in the program. ! 229: ! 230: Another source of incompatibility is that VAX-C assumes that: ! 231: ! 232: #include "foobar" ! 233: ! 234: is actually asking for the file `foobar.h'. GNU CC does not make this ! 235: assumption, and instead takes what you ask for literally; it tries to ! 236: read the file `foobar'. The best way to avoid this problem is to ! 237: always specify the desired file extension in your include directives. ! 238: ! 239: GNU CC for VMS is distributed with a set of include files that is ! 240: sufficient to compile most general purpose programs. Even though the ! 241: GNU CC distribution does not contain header files to define constants ! 242: and structures for some VMS system-specific functions, there is no ! 243: reason why you cannot use GNU CC with any of these functions. You ! 244: first may have to generate or create header files, either by using the ! 245: public domain utility `UNSDL' (which can be found on a DECUS tape), or ! 246: by extracting the relevant modules from one of the system macro ! 247: libraries, and using an editor to construct a C header file. ! 248: ! 249: ! 250: File: gcc.info, Node: Global Declarations, Next: VMS Misc, Prev: Include Files and VMS, Up: VMS ! 251: ! 252: Global Declarations and VMS ! 253: =========================== ! 254: ! 255: GNU CC does not provide the `globalref', `globaldef' and ! 256: `globalvalue' keywords of VAX-C. You can get the same effect with an ! 257: obscure feature of GAS, the GNU assembler. (This requires GAS version ! 258: 1.39 or later.) The following macros allow you to use this feature in ! 259: a fairly natural way: ! 260: ! 261: #ifdef __GNUC__ ! 262: #define GLOBALREF(NAME) \ ! 263: NAME asm("_$$PsectAttributes_GLOBALSYMBOL$$" #NAME ) ! 264: #define GLOBALDEF(NAME,VALUE) \ ! 265: NAME asm("_$$PsectAttributes_GLOBALSYMBOL$$" #NAME ) = VALUE ! 266: #define GLOBALVALUEREF(NAME) \ ! 267: const NAME [1] asm("_$$PsectAttributes_GLOBALVALUE$$" #NAME ) ! 268: #define GLOBALVALUEDEF(NAME,VALUE) \ ! 269: const NAME [1] asm("_$$PsectAttributes_GLOBALVALUE$$" #NAME ) = {VALUE} ! 270: #else ! 271: #define GLOBALREF(NAME) globalref NAME ! 272: #define GLOBALDEF(NAME,VALUE) globaldef NAME = VALUE ! 273: #define GLOBALVALUEDEF(NAME,VALUE) globalvalue NAME = VALUE ! 274: #define GLOBALVALUEREF(NAME) globalvalue NAME ! 275: #endif ! 276: ! 277: (The `_$$PsectAttributes_GLOBALSYMBOL' prefix at the start of the name ! 278: is removed by the assembler, after it has modified the attributes of ! 279: the symbol). These macros are provided in the VMS binaries ! 280: distribution in a header file `GNU_HACKS.H'. An example of the usage ! 281: is: ! 282: ! 283: int GLOBALREF (ijk); ! 284: int GLOBALDEF (jkl, 0); ! 285: ! 286: The macros `GLOBALREF' and `GLOBALDEF' cannot be used ! 287: straightforwardly for arrays, since there is no way to insert the array ! 288: dimension into the declaration at the right place. However, you can ! 289: declare an array with these macros if you first define a typedef for ! 290: the array type, like this: ! 291: ! 292: typedef int intvector[10]; ! 293: intvector GLOBALREF (foo); ! 294: ! 295: Array and structure initializers will also break the macros; you can ! 296: define the initializer to be a macro of its own, or you can expand the ! 297: `GLOBALDEF' macro by hand. You may find a case where you wish to use ! 298: the `GLOBALDEF' macro with a large array, but you are not interested ! 299: in explicitly initializing each element of the array. In such cases ! 300: you can use an initializer like: `{0,}', which will initialize the ! 301: entire array to `0'. ! 302: ! 303: A shortcoming of this implementation is that a variable declared ! 304: with `GLOBALVALUEREF' or `GLOBALVALUEDEF' is always an array. For ! 305: example, the declaration: ! 306: ! 307: int GLOBALVALUEREF(ijk); ! 308: ! 309: declares the variable `ijk' as an array of type `int [1]'. This is ! 310: done because a globalvalue is actually a constant; its "value" is what ! 311: the linker would normally consider an address. That is not how an ! 312: integer value works in C, but it is how an array works. So treating ! 313: the symbol as an array name gives consistent results--with the ! 314: exception that the value seems to have the wrong type. *Don't try to ! 315: access an element of the array.* It doesn't have any elements. The ! 316: array "address" may not be the address of actual storage. ! 317: ! 318: The fact that the symbol is an array may lead to warnings where the ! 319: variable is used. Insert type casts to avoid the warnings. Here is an ! 320: example; it takes advantage of the ANSI C feature allowing macros that ! 321: expand to use the same name as the macro itself. ! 322: ! 323: int GLOBALVALUEREF (ss$_normal); ! 324: int GLOBALVALUEDEF (xyzzy,123); ! 325: #ifdef __GNUC__ ! 326: #define ss$_normal ((int) ss$_normal) ! 327: #define xyzzy ((int) xyzzy) ! 328: #endif ! 329: ! 330: Don't use `globaldef' or `globalref' with a variable whose type is ! 331: an enumeration type; this is not implemented. Instead, make the ! 332: variable an integer, and use a `globalvaluedef' for each of the ! 333: enumeration values. An example of this would be: ! 334: ! 335: #ifdef __GNUC__ ! 336: int GLOBALDEF (color, 0); ! 337: int GLOBALVALUEDEF (RED, 0); ! 338: int GLOBALVALUEDEF (BLUE, 1); ! 339: int GLOBALVALUEDEF (GREEN, 3); ! 340: #else ! 341: enum globaldef color {RED, BLUE, GREEN = 3}; ! 342: #endif ! 343: ! 344: ! 345: File: gcc.info, Node: VMS Misc, Prev: Global Declarations, Up: VMS ! 346: ! 347: Other VMS Issues ! 348: ================ ! 349: ! 350: GNU CC automatically arranges for `main' to return 1 by default if ! 351: you fail to specify an explicit return value. This will be interpreted ! 352: by VMS as a status code indicating a normal successful completion. ! 353: Version 1 of GNU CC did not provide this default. ! 354: ! 355: GNU CC on VMS works only with the GNU assembler, GAS. You need ! 356: version 1.37 or later of GAS in order to produce value debugging ! 357: information for the VMS debugger. Use the ordinary VMS linker with ! 358: the object files produced by GAS. ! 359: ! 360: Under previous versions of GNU CC, the generated code would ! 361: occasionally give strange results when linked to the sharable ! 362: `VAXCRTL' library. Now this should work. ! 363: ! 364: A caveat for use of `const' global variables: the `const' modifier ! 365: must be specified in every external declaration of the variable in all ! 366: of the source files that use that variable. Otherwise the linker will ! 367: issue warnings about conflicting attributes for the variable. Your ! 368: program will still work despite the warnings, but the variable will be ! 369: placed in writable storage. ! 370: ! 371: The VMS linker does not distinguish between upper and lower case ! 372: letters in function and variable names. However, usual practice in C ! 373: is to distinguish case. Normally GNU CC (by means of the assembler ! 374: GAS) implements usual C behavior by augmenting each name that is not ! 375: all lower-case. A name is augmented by truncating it to at most 23 ! 376: characters and then adding more characters at the end which encode the ! 377: case pattern the rest. ! 378: ! 379: Name augmentation yields bad results for programs that use ! 380: precompiled libraries (such as Xlib) which were generated by another ! 381: compiler. You can use the compiler option `/NOCASE_HACK' to inhibit ! 382: augmentation; it makes external C functions and variables ! 383: case-independent as is usual on VMS. Alternatively, you could write ! 384: all references to the functions and variables in such libraries using ! 385: lower case; this will work on VMS, but is not portable to other ! 386: systems. ! 387: ! 388: Function and variable names are handled somewhat differently with ! 389: GNU C++. The GNU C++ compiler performs "name mangling" on function ! 390: names, which means that it adds information to the function name to ! 391: describe the data types of the arguments that the function takes. One ! 392: result of this is that the name of a function can become very long. ! 393: Since the VMS linker only recognizes the first 31 characters in a name, ! 394: special action is taken to ensure that each function and variable has a ! 395: unique name that can be represented in 31 characters. ! 396: ! 397: If the name (plus a name augmentation, if required) is less than 32 ! 398: characters in length, then no special action is performed. If the name ! 399: is longer than 31 characters, the assembler (GAS) will generate a hash ! 400: string based upon the function name, truncate the function name to 23 ! 401: characters, and append the hash string to the truncated name. If the ! 402: `/VERBOSE' compiler option is used, the assembler will print both the ! 403: full and truncated names of each symbol that is truncated. ! 404: ! 405: The `/NOCASE_HACK' compiler option should not be used when you are ! 406: compiling programs that use libg++. libg++ has several instances of ! 407: objects (i.e. `Filebuf' and `filebuf') which become indistinguishable ! 408: in a case-insensitive environment. This leads to cases where you need ! 409: to inhibit augmentation selectively (if you were using libg++ and Xlib ! 410: in the same program, for example). There is no special feature for ! 411: doing this, but you can get the result by defining a macro for each ! 412: mixed case symbol for which you wish to inhibit augmentation. The ! 413: macro should expand into the lower case equivalent of itself. For ! 414: example: ! 415: ! 416: #define StuDlyCapS studlycaps ! 417: ! 418: These macro definitions can be placed in a header file to minimize ! 419: the number of changes to your source code. ! 420: ! 421: ! 422: File: gcc.info, Node: Portability, Next: Interface, Prev: VMS, Up: Top ! 423: ! 424: GNU CC and Portability ! 425: ********************** ! 426: ! 427: The main goal of GNU CC was to make a good, fast compiler for ! 428: machines in the class that the GNU system aims to run on: 32-bit ! 429: machines that address 8-bit bytes and have several general registers. ! 430: Elegance, theoretical power and simplicity are only secondary. ! 431: ! 432: GNU CC gets most of the information about the target machine from a ! 433: machine description which gives an algebraic formula for each of the ! 434: machine's instructions. This is a very clean way to describe the ! 435: target. But when the compiler needs information that is difficult to ! 436: express in this fashion, I have not hesitated to define an ad-hoc ! 437: parameter to the machine description. The purpose of portability is ! 438: to reduce the total work needed on the compiler; it was not of ! 439: interest for its own sake. ! 440: ! 441: GNU CC does not contain machine dependent code, but it does contain ! 442: code that depends on machine parameters such as endianness (whether ! 443: the most significant byte has the highest or lowest address of the ! 444: bytes in a word) and the availability of autoincrement addressing. In ! 445: the RTL-generation pass, it is often necessary to have multiple ! 446: strategies for generating code for a particular kind of syntax tree, ! 447: strategies that are usable for different combinations of parameters. ! 448: Often I have not tried to address all possible cases, but only the ! 449: common ones or only the ones that I have encountered. As a result, a ! 450: new target may require additional strategies. You will know if this ! 451: happens because the compiler will call `abort'. Fortunately, the new ! 452: strategies can be added in a machine-independent fashion, and will ! 453: affect only the target machines that need them. ! 454: ! 455: ! 456: File: gcc.info, Node: Interface, Next: Passes, Prev: Portability, Up: Top ! 457: ! 458: Interfacing to GNU CC Output ! 459: **************************** ! 460: ! 461: GNU CC is normally configured to use the same function calling ! 462: convention normally in use on the target system. This is done with the ! 463: machine-description macros described (*note Machine Macros::.). ! 464: ! 465: However, returning of structure and union values is done ! 466: differently on some target machines. As a result, functions compiled ! 467: with PCC returning such types cannot be called from code compiled with ! 468: GNU CC, and vice versa. This does not cause trouble often because few ! 469: Unix library routines return structures or unions. ! 470: ! 471: GNU CC code returns structures and unions that are 1, 2, 4 or 8 ! 472: bytes long in the same registers used for `int' or `double' return ! 473: values. (GNU CC typically allocates variables of such types in ! 474: registers also.) Structures and unions of other sizes are returned by ! 475: storing them into an address passed by the caller (usually in a ! 476: register). The machine-description macros `STRUCT_VALUE' and ! 477: `STRUCT_INCOMING_VALUE' tell GNU CC where to pass this address. ! 478: ! 479: By contrast, PCC on most target machines returns structures and ! 480: unions of any size by copying the data into an area of static storage, ! 481: and then returning the address of that storage as if it were a pointer ! 482: value. The caller must copy the data from that memory area to the ! 483: place where the value is wanted. This is slower than the method used ! 484: by GNU CC, and fails to be reentrant. ! 485: ! 486: On some target machines, such as RISC machines and the 80386, the ! 487: standard system convention is to pass to the subroutine the address of ! 488: where to return the value. On these machines, GNU CC has been ! 489: configured to be compatible with the standard compiler, when this ! 490: method is used. It may not be compatible for structures of 1, 2, 4 or ! 491: 8 bytes. ! 492: ! 493: GNU CC uses the system's standard convention for passing arguments. ! 494: On some machines, the first few arguments are passed in registers; in ! 495: others, all are passed on the stack. It would be possible to use ! 496: registers for argument passing on any machine, and this would probably ! 497: result in a significant speedup. But the result would be complete ! 498: incompatibility with code that follows the standard convention. So ! 499: this change is practical only if you are switching to GNU CC as the ! 500: sole C compiler for the system. We may implement register argument ! 501: passing on certain machines once we have a complete GNU system so that ! 502: we can compile the libraries with GNU CC. ! 503: ! 504: On some machines (particularly the Sparc), certain types of ! 505: arguments are passed "by invisible reference". This means that the ! 506: value is stored in memory, and the address of the memory location is ! 507: passed to the subroutine. ! 508: ! 509: If you use `longjmp', beware of automatic variables. ANSI C says ! 510: that automatic variables that are not declared `volatile' have ! 511: undefined values after a `longjmp'. And this is all GNU CC promises ! 512: to do, because it is very difficult to restore register variables ! 513: correctly, and one of GNU CC's features is that it can put variables ! 514: in registers without your asking it to. ! 515: ! 516: If you want a variable to be unaltered by `longjmp', and you don't ! 517: want to write `volatile' because old C compilers don't accept it, just ! 518: take the address of the variable. If a variable's address is ever ! 519: taken, even if just to compute it and ignore it, then the variable ! 520: cannot go in a register: ! 521: ! 522: { ! 523: int careful; ! 524: &careful; ! 525: ... ! 526: } ! 527: ! 528: Code compiled with GNU CC may call certain library routines. Most ! 529: of them handle arithmetic for which there are no instructions. This ! 530: includes multiply and divide on some machines, and floating point ! 531: operations on any machine for which floating point support is disabled ! 532: with `-msoft-float'. Some standard parts of the C library, such as ! 533: `bcopy' or `memcpy', are also called automatically. The usual ! 534: function call interface is used for calling the library routines. ! 535: ! 536: These library routines should be defined in the library `libgcc.a', ! 537: which GNU CC automatically searches whenever it links a program. On ! 538: machines that have multiply and divide instructions, if hardware ! 539: floating point is in use, normally `libgcc.a' is not needed, but it is ! 540: searched just in case. ! 541: ! 542: Each arithmetic function is defined in `libgcc1.c' to use the ! 543: corresponding C arithmetic operator. As long as the file is compiled ! 544: with another C compiler, which supports all the C arithmetic operators, ! 545: this file will work portably. However, `libgcc1.c' does not work if ! 546: compiled with GNU CC, because each arithmetic function would compile ! 547: into a call to itself! ! 548: ! 549: ! 550: File: gcc.info, Node: Passes, Next: RTL, Prev: Interface, Up: Top ! 551: ! 552: Passes and Files of the Compiler ! 553: ******************************** ! 554: ! 555: The overall control structure of the compiler is in `toplev.c'. ! 556: This file is responsible for initialization, decoding arguments, ! 557: opening and closing files, and sequencing the passes. ! 558: ! 559: The parsing pass is invoked only once, to parse the entire input. ! 560: The RTL intermediate code for a function is generated as the function ! 561: is parsed, a statement at a time. Each statement is read in as a ! 562: syntax tree and then converted to RTL; then the storage for the tree ! 563: for the statement is reclaimed. Storage for types (and the ! 564: expressions for their sizes), declarations, and a representation of ! 565: the binding contours and how they nest, remain until the function is ! 566: finished being compiled; these are all needed to output the debugging ! 567: information. ! 568: ! 569: Each time the parsing pass reads a complete function definition or ! 570: top-level declaration, it calls the function `rest_of_compilation' or ! 571: `rest_of_decl_compilation' in `toplev.c', which are responsible for ! 572: all further processing necessary, ending with output of the assembler ! 573: language. All other compiler passes run, in sequence, within ! 574: `rest_of_compilation'. When that function returns from compiling a ! 575: function definition, the storage used for that function definition's ! 576: compilation is entirely freed, unless it is an inline function (*note ! 577: Inline::.). ! 578: ! 579: Here is a list of all the passes of the compiler and their source ! 580: files. Also included is a description of where debugging dumps can be ! 581: requested with `-d' options. ! 582: ! 583: * Parsing. This pass reads the entire text of a function ! 584: definition, constructing partial syntax trees. This and RTL ! 585: generation are no longer truly separate passes (formerly they ! 586: were), but it is easier to think of them as separate. ! 587: ! 588: The tree representation does not entirely follow C syntax, ! 589: because it is intended to support other languages as well. ! 590: ! 591: Language-specific data type analysis is also done in this pass, ! 592: and every tree node that represents an expression has a data type ! 593: attached. Variables are represented as declaration nodes. ! 594: ! 595: Constant folding and some arithmetic simplifications are also done ! 596: during this pass. ! 597: ! 598: The language-independent source files for parsing are ! 599: `stor-layout.c', `fold-const.c', and `tree.c'. There are also ! 600: header files `tree.h' and `tree.def' which define the format of ! 601: the tree representation. ! 602: ! 603: The source files for parsing C are `c-parse.y', `c-decl.c', ! 604: `c-typeck.c', `c-convert.c', `c-lang.c', and `c-aux-info.c' along ! 605: with header files `c-lex.h', and `c-tree.h'. ! 606: ! 607: The source files for parsing C++ are `cp-parse.y', `cp-class.c', ! 608: `cp-cvt.c', ! 609: `cp-decl.c', `cp-decl.c', `cp-decl2.c', `cp-dem.c', ! 610: `cp-except.c', ! 611: `cp-expr.c', `cp-init.c', `cp-lex.c', `cp-method.c', ! 612: `cp-ptree.c', ! 613: `cp-search.c', `cp-tree.c', `cp-type2.c', and `cp-typeck.c', ! 614: along with header files `cp-tree.def', `cp-tree.h', and ! 615: `cp-decl.h'. ! 616: ! 617: The special source files for parsing Objective C are ! 618: `objc-parse.y', `objc-actions.c', `objc-tree.def', and ! 619: `objc-actions.h'. Certain C-specific files are used for this as ! 620: well. ! 621: ! 622: The file `c-common.c' is also used for all of the above languages. ! 623: ! 624: * RTL generation. This is the conversion of syntax tree into RTL ! 625: code. It is actually done statement-by-statement during parsing, ! 626: but for most purposes it can be thought of as a separate pass. ! 627: ! 628: This is where the bulk of target-parameter-dependent code is ! 629: found, since often it is necessary for strategies to apply only ! 630: when certain standard kinds of instructions are available. The ! 631: purpose of named instruction patterns is to provide this ! 632: information to the RTL generation pass. ! 633: ! 634: Optimization is done in this pass for `if'-conditions that are ! 635: comparisons, boolean operations or conditional expressions. Tail ! 636: recursion is detected at this time also. Decisions are made ! 637: about how best to arrange loops and how to output `switch' ! 638: statements. ! 639: ! 640: The source files for RTL generation include `stmt.c', ! 641: `function.c', `expr.c', `calls.c', `explow.c', `expmed.c', ! 642: `optabs.c' and `emit-rtl.c'. Also, the file `insn-emit.c', ! 643: generated from the machine description by the program `genemit', ! 644: is used in this pass. The header file `expr.h' is used for ! 645: communication within this pass. ! 646: ! 647: The header files `insn-flags.h' and `insn-codes.h', generated ! 648: from the machine description by the programs `genflags' and ! 649: `gencodes', tell this pass which standard names are available for ! 650: use and which patterns correspond to them. ! 651: ! 652: Aside from debugging information output, none of the following ! 653: passes refers to the tree structure representation of the ! 654: function (only part of which is saved). ! 655: ! 656: The decision of whether the function can and should be expanded ! 657: inline in its subsequent callers is made at the end of rtl ! 658: generation. The function must meet certain criteria, currently ! 659: related to the size of the function and the types and number of ! 660: parameters it has. Note that this function may contain loops, ! 661: recursive calls to itself (tail-recursive functions can be ! 662: inlined!), gotos, in short, all constructs supported by GNU CC. ! 663: The file `integrate.c' contains the code to save a function's rtl ! 664: for later inlining and to inline that rtl when the function is ! 665: called. The header file `integrate.h' is also used for this ! 666: purpose. ! 667: ! 668: The option `-dr' causes a debugging dump of the RTL code after ! 669: this pass. This dump file's name is made by appending `.rtl' to ! 670: the input file name. ! 671: ! 672: * Jump optimization. This pass simplifies jumps to the following ! 673: instruction, jumps across jumps, and jumps to jumps. It deletes ! 674: unreferenced labels and unreachable code, except that unreachable ! 675: code that contains a loop is not recognized as unreachable in ! 676: this pass. (Such loops are deleted later in the basic block ! 677: analysis.) It also converts some code originally written with ! 678: jumps into sequences of instructions that directly set values ! 679: from the results of comparisons, if the machine has such ! 680: instructions. ! 681: ! 682: Jump optimization is performed two or three times. The first ! 683: time is immediately following RTL generation. The second time is ! 684: after CSE, but only if CSE says repeated jump optimization is ! 685: needed. The last time is right before the final pass. That ! 686: time, cross-jumping and deletion of no-op move instructions are ! 687: done together with the optimizations described above. ! 688: ! 689: The source file of this pass is `jump.c'. ! 690: ! 691: The option `-dj' causes a debugging dump of the RTL code after ! 692: this pass is run for the first time. This dump file's name is ! 693: made by appending `.jump' to the input file name. ! 694: ! 695: * Register scan. This pass finds the first and last use of each ! 696: register, as a guide for common subexpression elimination. Its ! 697: source is in `regclass.c'. ! 698: ! 699: * Jump threading. This pass detects a condition jump that branches ! 700: to an identical or inverse test. Such jumps can be `threaded' ! 701: through the second conditional test. The source code for this ! 702: pass is in `jump.c'. This optimization is only performed if ! 703: `-fthread-jumps' is enabled. ! 704: ! 705: * Common subexpression elimination. This pass also does constant ! 706: propagation. Its source file is `cse.c'. If constant ! 707: propagation causes conditional jumps to become unconditional or to ! 708: become no-ops, jump optimization is run again when CSE is ! 709: finished. ! 710: ! 711: The option `-ds' causes a debugging dump of the RTL code after ! 712: this pass. This dump file's name is made by appending `.cse' to ! 713: the input file name. ! 714: ! 715: * Loop optimization. This pass moves constant expressions out of ! 716: loops, and optionally does strength-reduction and loop unrolling ! 717: as well. Its source files are `loop.c' and `unroll.c', plus the ! 718: header `loop.h' used for communication between them. Loop ! 719: unrolling uses some functions in `integrate.c' and the header ! 720: `integrate.h'. ! 721: ! 722: The option `-dL' causes a debugging dump of the RTL code after ! 723: this pass. This dump file's name is made by appending `.loop' to ! 724: the input file name. ! 725: ! 726: * If `-frerun-cse-after-loop' was enabled, a second common ! 727: subexpression elimination pass is performed after the loop ! 728: optimization pass. Jump threading is also done again at this ! 729: time if it was specified. ! 730: ! 731: The option `-dt' causes a debugging dump of the RTL code after ! 732: this pass. This dump file's name is made by appending `.cse2' to ! 733: the input file name. ! 734: ! 735: * Stupid register allocation is performed at this point in a ! 736: nonoptimizing compilation. It does a little data flow analysis as ! 737: well. When stupid register allocation is in use, the next pass ! 738: executed is the reloading pass; the others in between are skipped. ! 739: The source file is `stupid.c'. ! 740: ! 741: * Data flow analysis (`flow.c'). This pass divides the program ! 742: into basic blocks (and in the process deletes unreachable loops); ! 743: then it computes which pseudo-registers are live at each point in ! 744: the program, and makes the first instruction that uses a value ! 745: point at the instruction that computed the value. ! 746: ! 747: This pass also deletes computations whose results are never used, ! 748: and combines memory references with add or subtract instructions ! 749: to make autoincrement or autodecrement addressing. ! 750: ! 751: The option `-df' causes a debugging dump of the RTL code after ! 752: this pass. This dump file's name is made by appending `.flow' to ! 753: the input file name. If stupid register allocation is in use, ! 754: this dump file reflects the full results of such allocation. ! 755: ! 756: * Instruction combination (`combine.c'). This pass attempts to ! 757: combine groups of two or three instructions that are related by ! 758: data flow into single instructions. It combines the RTL ! 759: expressions for the instructions by substitution, simplifies the ! 760: result using algebra, and then attempts to match the result ! 761: against the machine description. ! 762: ! 763: The option `-dc' causes a debugging dump of the RTL code after ! 764: this pass. This dump file's name is made by appending `.combine' ! 765: to the input file name. ! 766: ! 767: * Instruction scheduling (`sched.c'). This pass looks for ! 768: instructions whose output will not be available by the time that ! 769: it is used in subsequent instructions. (Memory loads and ! 770: floating point instructions often have this behavior on RISC ! 771: machines). It re-orders instructions within a basic block to try ! 772: to separate the definition and use of items that otherwise would ! 773: cause pipeline stalls. ! 774: ! 775: Instruction scheduling is performed twice. The first time is ! 776: immediately after instruction combination and the second is ! 777: immediately after reload. ! 778: ! 779: The option `-dS' causes a debugging dump of the RTL code after ! 780: this pass is run for the first time. The dump file's name is ! 781: made by appending `.sched' to the input file name. ! 782: ! 783: * Register class preferencing. The RTL code is scanned to find out ! 784: which register class is best for each pseudo register. The source ! 785: file is `regclass.c'. ! 786: ! 787: * Local register allocation (`local-alloc.c'). This pass allocates ! 788: hard registers to pseudo registers that are used only within one ! 789: basic block. Because the basic block is linear, it can use fast ! 790: and powerful techniques to do a very good job. ! 791: ! 792: The option `-dl' causes a debugging dump of the RTL code after ! 793: this pass. This dump file's name is made by appending `.lreg' to ! 794: the input file name. ! 795: ! 796: * Global register allocation (`global-alloc.c'). This pass ! 797: allocates hard registers for the remaining pseudo registers (those ! 798: whose life spans are not contained in one basic block). ! 799: ! 800: * Reloading. This pass renumbers pseudo registers with the hardware ! 801: registers numbers they were allocated. Pseudo registers that did ! 802: not get hard registers are replaced with stack slots. Then it ! 803: finds instructions that are invalid because a value has failed to ! 804: end up in a register, or has ended up in a register of the wrong ! 805: kind. It fixes up these instructions by reloading the ! 806: problematical values temporarily into registers. Additional ! 807: instructions are generated to do the copying. ! 808: ! 809: The reload pass also optionally eliminates the frame pointer and ! 810: inserts instructions to save and restore call-clobbered registers ! 811: around calls. ! 812: ! 813: Source files are `reload.c' and `reload1.c', plus the header ! 814: `reload.h' used for communication between them. ! 815: ! 816: The option `-dg' causes a debugging dump of the RTL code after ! 817: this pass. This dump file's name is made by appending `.greg' to ! 818: the input file name. ! 819: ! 820: * Instruction scheduling is repeated here to try to avoid pipeline ! 821: stalls due to memory loads generated for spilled pseudo registers. ! 822: ! 823: The option `-dR' causes a debugging dump of the RTL code after ! 824: this pass. This dump file's name is made by appending `.sched2' ! 825: to the input file name. ! 826: ! 827: * Jump optimization is repeated, this time including cross-jumping ! 828: and deletion of no-op move instructions. ! 829: ! 830: The option `-dJ' causes a debugging dump of the RTL code after ! 831: this pass. This dump file's name is made by appending `.jump2' ! 832: to the input file name. ! 833: ! 834: * Delayed branch scheduling. This optional pass attempts to find ! 835: instructions that can go into the delay slots of other ! 836: instructions, usually jumps and calls. The source file name is ! 837: `reorg.c'. ! 838: ! 839: The option `-dd' causes a debugging dump of the RTL code after ! 840: this pass. This dump file's name is made by appending `.dbr' to ! 841: the input file name. ! 842: ! 843: * Conversion from usage of some hard registers to usage of a ! 844: register stack may be done at this point. Currently, this is ! 845: supported only for the floating-point registers of the Intel ! 846: 80387 coprocessor. The source file name is `reg-stack.c'. ! 847: ! 848: The options `-dk' causes a debugging dump of the RTL code after ! 849: this pass. This dump file's name is made by appending `.stack' ! 850: to the input file name. ! 851: ! 852: * Final. This pass outputs the assembler code for the function. ! 853: It is also responsible for identifying spurious test and compare ! 854: instructions. Machine-specific peephole optimizations are ! 855: performed at the same time. The function entry and exit ! 856: sequences are generated directly as assembler code in this pass; ! 857: they never exist as RTL. ! 858: ! 859: The source files are `final.c' plus `insn-output.c'; the latter ! 860: is generated automatically from the machine description by the ! 861: tool `genoutput'. The header file `conditions.h' is used for ! 862: communication between these files. ! 863: ! 864: * Debugging information output. This is run after final because it ! 865: must output the stack slot offsets for pseudo registers that did ! 866: not get hard registers. Source files are `dbxout.c' for DBX ! 867: symbol table format, `sdbout.c' for SDB symbol table format, and ! 868: `dwarfout.c' for DWARF symbol table format. ! 869: ! 870: Some additional files are used by all or many passes: ! 871: ! 872: * Every pass uses `machmode.def' and `machmode.h' which define the ! 873: machine modes. ! 874: ! 875: * Several passes use `real.h', which defines the default ! 876: representation of floating point constants and how to operate on ! 877: them. ! 878: ! 879: * All the passes that work with RTL use the header files `rtl.h' ! 880: and `rtl.def', and subroutines in file `rtl.c'. The tools `gen*' ! 881: also use these files to read and work with the machine ! 882: description RTL. ! 883: ! 884: * Several passes refer to the header file `insn-config.h' which ! 885: contains a few parameters (C macro definitions) generated ! 886: automatically from the machine description RTL by the tool ! 887: `genconfig'. ! 888: ! 889: * Several passes use the instruction recognizer, which consists of ! 890: `recog.c' and `recog.h', plus the files `insn-recog.c' and ! 891: `insn-extract.c' that are generated automatically from the ! 892: machine description by the tools `genrecog' and `genextract'. ! 893: ! 894: * Several passes use the header files `regs.h' which defines the ! 895: information recorded about pseudo register usage, and ! 896: `basic-block.h' which defines the information recorded about ! 897: basic blocks. ! 898: ! 899: * `hard-reg-set.h' defines the type `HARD_REG_SET', a bit-vector ! 900: with a bit for each hard register, and some macros to manipulate ! 901: it. This type is just `int' if the machine has few enough hard ! 902: registers; otherwise it is an array of `int' and some of the ! 903: macros expand into loops. ! 904: ! 905: * Several passes use instruction attributes. A definition of the ! 906: attributes defined for a particular machine is in file ! 907: `insn-attr.h', which is generated from the machine description by ! 908: the program `genattr'. The file `insn-attrtab.c' contains ! 909: subroutines to obtain the attribute values for insns. It is ! 910: generated from the machine description by the program ! 911: `genattrtab'. ! 912: ! 913: ! 914: File: gcc.info, Node: RTL, Next: Machine Desc, Prev: Passes, Up: Top ! 915: ! 916: RTL Representation ! 917: ****************** ! 918: ! 919: Most of the work of the compiler is done on an intermediate ! 920: representation called register transfer language. In this language, ! 921: the instructions to be output are described, pretty much one by one, ! 922: in an algebraic form that describes what the instruction does. ! 923: ! 924: RTL is inspired by Lisp lists. It has both an internal form, made ! 925: up of structures that point at other structures, and a textual form ! 926: that is used in the machine description and in printed debugging ! 927: dumps. The textual form uses nested parentheses to indicate the ! 928: pointers in the internal form. ! 929: ! 930: * Menu: ! 931: ! 932: * RTL Objects:: Expressions vs vectors vs strings vs integers. ! 933: * Accessors:: Macros to access expression operands or vector elts. ! 934: * Flags:: Other flags in an RTL expression. ! 935: * Machine Modes:: Describing the size and format of a datum. ! 936: * Constants:: Expressions with constant values. ! 937: * Regs and Memory:: Expressions representing register contents or memory. ! 938: * Arithmetic:: Expressions representing arithmetic on other expressions. ! 939: * Comparisons:: Expressions representing comparison of expressions. ! 940: * Bit Fields:: Expressions representing bit-fields in memory or reg. ! 941: * Conversions:: Extending, truncating, floating or fixing. ! 942: * RTL Declarations:: Declaring volatility, constancy, etc. ! 943: * Side Effects:: Expressions for storing in registers, etc. ! 944: * Incdec:: Embedded side-effects for autoincrement addressing. ! 945: * Assembler:: Representing `asm' with operands. ! 946: * Insns:: Expression types for entire insns. ! 947: * Calls:: RTL representation of function call insns. ! 948: * Sharing:: Some expressions are unique; others *must* be copied. ! 949: ! 950: ! 951: File: gcc.info, Node: RTL Objects, Next: Accessors, Prev: RTL, Up: RTL ! 952: ! 953: RTL Object Types ! 954: ================ ! 955: ! 956: RTL uses four kinds of objects: expressions, integers, strings and ! 957: vectors. Expressions are the most important ones. An RTL expression ! 958: ("RTX", for short) is a C structure, but it is usually referred to ! 959: with a pointer; a type that is given the typedef name `rtx'. ! 960: ! 961: An integer is simply an `int'; their written form uses decimal ! 962: digits. ! 963: ! 964: A string is a sequence of characters. In core it is represented as ! 965: a `char *' in usual C fashion, and it is written in C syntax as well. ! 966: However, strings in RTL may never be null. If you write an empty ! 967: string in a machine description, it is represented in core as a null ! 968: pointer rather than as a pointer to a null character. In certain ! 969: contexts, these null pointers instead of strings are valid. Within ! 970: RTL code, strings are most commonly found inside `symbol_ref' ! 971: expressions, but they appear in other contexts in the RTL expressions ! 972: that make up machine descriptions. ! 973: ! 974: A vector contains an arbitrary number of pointers to expressions. ! 975: The number of elements in the vector is explicitly present in the ! 976: vector. The written form of a vector consists of square brackets ! 977: (`[...]') surrounding the elements, in sequence and with whitespace ! 978: separating them. Vectors of length zero are not created; null ! 979: pointers are used instead. ! 980: ! 981: Expressions are classified by "expression codes" (also called RTX ! 982: codes). The expression code is a name defined in `rtl.def', which is ! 983: also (in upper case) a C enumeration constant. The possible expression ! 984: codes and their meanings are machine-independent. The code of an RTX ! 985: can be extracted with the macro `GET_CODE (X)' and altered with ! 986: `PUT_CODE (X, NEWCODE)'. ! 987: ! 988: The expression code determines how many operands the expression ! 989: contains, and what kinds of objects they are. In RTL, unlike Lisp, ! 990: you cannot tell by looking at an operand what kind of object it is. ! 991: Instead, you must know from its context--from the expression code of ! 992: the containing expression. For example, in an expression of code ! 993: `subreg', the first operand is to be regarded as an expression and the ! 994: second operand as an integer. In an expression of code `plus', there ! 995: are two operands, both of which are to be regarded as expressions. In ! 996: a `symbol_ref' expression, there is one operand, which is to be ! 997: regarded as a string. ! 998: ! 999: Expressions are written as parentheses containing the name of the ! 1000: expression type, its flags and machine mode if any, and then the ! 1001: operands of the expression (separated by spaces). ! 1002: ! 1003: Expression code names in the `md' file are written in lower case, ! 1004: but when they appear in C code they are written in upper case. In this ! 1005: manual, they are shown as follows: `const_int'. ! 1006: ! 1007: In a few contexts a null pointer is valid where an expression is ! 1008: normally wanted. The written form of this is `(nil)'. ! 1009: ! 1010:
This archive runs on limited infrastructure. Preserving old code on modern bandwidth. Automated agents are requested to crawl responsibly.