--- gcc/internals-2 2018/04/24 16:37:52 1.1.1.1 +++ gcc/internals-2 2018/04/24 16:39:26 1.1.1.3 @@ -1,1349 +1,945 @@ -Info file internals, produced by texinfo-format-buffer -*-Text-*- -from file internals.texinfo + +File: internals, Node: Extensions, Next: Bugs, Prev: Incompatibilities, Up: Top + +GNU Extensions to the C Language +******************************** + +GNU C provides several language features not found in ANSI standard C. +(The `-pedantic' option directs GNU CC to print a warning message if any of +these features is used.) To test for the availability of these features in +conditional compilation, check for a predefined macro `__GNUC__', which is +always defined under GNU CC. + +* Menu: + +* Statement Exprs:: Putting statements and declarations inside expressions. +* Naming Types:: Giving a name to the type of some expression. +* Typeof:: `typeof': referring to the type of an expression. +* Lvalues:: Using `?:', `,' and casts in lvalues. +* Conditionals:: Omitting the middle operand of a `?:' expression. +* Zero-Length:: Zero-length arrays. +* Variable-Length:: Arrays whose length is computed at run time. +* Subscripting:: Any array can be subscripted, even if not an lvalue. +* Pointer Arith:: Arithmetic on `void'-pointers and function pointers. +* Constructors:: Constructor expressions give structures, unions + or arrays as values. +* Dollar Signs:: Dollar sign is allowed in identifiers. +* Alignment:: Inquiring about the alignment of a type or variable. +* Inline:: Defining inline functions (as fast as macros). +* Extended Asm:: Assembler instructions with C expressions as operands. + (With them you can define ``built-in'' functions.) +* Asm Labels:: Specifying the assembler name to use for a C symbol. + + + +File: internals, Node: Statement Exprs, Next: Naming Types, Prev: Extensions, Up: Extensions + +Statements and Declarations inside of Expressions +================================================= + +A compound statement in parentheses may appear inside an expression in GNU +C. This allows you to declare variables within an expression. For example: + + ({ int y = foo (); int z; + if (y > 0) z = y; + else z = - y; + z; }) + +is a valid (though slightly more complex than necessary) expression for the +absolute value of `foo ()'. + +This feature is especially useful in making macro definitions ``safe'' (so +that they evaluate each operand exactly once). For example, the +``maximum'' function is commonly defined as a macro in standard C as follows: + + #define max(a,b) ((a) > (b) ? (a) : (b)) + +But this definition computes either A or B twice, with bad results if the +operand has side effects. In GNU C, if you know the type of the operands +(here let's assume `int'), you can define the macro safely as follows: + + #define maxint(a,b) \ + ({int _a = (a), _b = (b); _a > _b ? _a : _b; }) + +Embedded statements are not allowed in constant expressions, such as the +value of an enumeration constant, the width of a bit field, or the initial +value of a static variable. + +If you don't know the type of the operand, you can still do this, but you +must use `typeof' (*Note Typeof::.) or type naming (*Note Naming Types::.). + + +File: internals, Node: Naming Types, Next: Typeof, Prev: Statement Exprs, Up: Extensions + +Naming an Expression's Type +=========================== + +You can give a name to the type of an expression using a `typedef' +declaration with an initializer. Here is how to define NAME as a type name +for the type of EXP: + + typedef NAME = EXP; + +This is useful in conjunction with the statements-within-expressions +feature. Here is how the two together can be used to define a safe +``maximum'' macro that operates on any arithmetic type: + + #define max(a,b) \ + ({typedef _ta = (a), _tb = (b); \ + _ta _a = (a); _tb _b = (b); \ + _a > _b ? _a : _b; }) + +The reason for using names that start with underscores for the local +variables is to avoid conflicts with variable names that occur within the +expressions that are substituted for `a' and `b'. Eventually we hope to +design a new form of declaration syntax that allows you to declare +variables whose scopes start only after their initializers; this will be a +more reliable way to prevent such conflicts. + + +File: internals, Node: Typeof, Next: Lvalues, Prev: Naming Types, Up: Extensions + +Referring to a Type with `typeof' +================================= + +Another way to refer to the type of an expression is with `typeof'. The +syntax of using of this keyword looks like `sizeof', but the construct acts +semantically like a type name defined with `typedef'. + +There are two ways of writing the argument to `typeof': with an expression +or with a type. Here is an example with an expression: + + typeof (x[0](1)) + +This assumes that `x' is an array of functions; the type described is that +of the values of the functions. + +Here is an example with a typename as the argument: + + typeof (int *) -This file documents the internals of the GNU compiler. +Here the type described is that of pointers to `int'. -Copyright (C) 1987 Richard M. Stallman. +A `typeof'-construct can be used anywhere a typedef name could be used. +For example, you can use it in a declaration, in a cast, or inside of +`sizeof' or `typeof'. -Permission is granted to make and distribute verbatim copies of -this manual provided the copyright notice and this permission notice -are preserved on all copies. + * This declares `y' with the type of what `x' points to. -Permission is granted to copy and distribute modified versions of this -manual under the conditions for verbatim copying, provided also that the -section entitled "GNU CC General Public License" is included exactly as -in the original, and provided that the entire resulting derived work is -distributed under the terms of a permission notice identical to this one. + typeof (*x) y; -Permission is granted to copy and distribute translations of this manual -into another language, under the above conditions for modified versions, -except that the section entitled "GNU CC General Public License" may be -included in a translation approved by the author instead of in the original -English. + * This declares `y' as an array of such values. + typeof (*x) y[4]; + * This declares `y' as an array of pointers to characters: + typeof (typeof (char *)[4]) y; + + It is equivalent to the following traditional C declaration: + + char *y[4]; + + To see the meaning of the declaration using `typeof', and why it might + be a useful way to write, let's rewrite it with these macros: + + #define pointer(T) typeof(T *) + #define array(T, N) typeof(T [N]) + + Now the declaration can be rewritten this way: + + array (pointer (char), 4) y; + + Thus, `array (pointer (char), 4)' is the type of arrays of 4 pointers + to `char'. + + +File: internals, Node: Lvalues, Next: Conditionals, Prev: Typeof, Up: Extensions + +Generalized Lvalues +=================== + +Compound expressions, conditional expressions and casts are allowed as +lvalues provided their operands are lvalues. This means that you can take +their addresses or store values into them. + +For example, a compound expression can be assigned, provided the last +expression in the sequence is an lvalue. These two expressions are +equivalent: + + (a, b) += 5 + a, (b += 5) + +Similarly, the address of the compound expression can be taken. These two +expressions are equivalent: + + &(a, b) + a, &b + +A conditional expression is a valid lvalue if its type is not void and the +true and false branches are both valid lvalues. For example, these two +expressions are equivalent: + + (a ? b : c) = 5 + (a ? b = 5 : (c = 5)) + +A cast is a valid lvalue if its operand is valid. Taking the address of +the cast is the same as taking the address without a cast, except for the +type of the result. For example, these two expressions are equivalent (but +the second may be valid when the type of `a' does not permit a cast to `int +*'). + + &(int *)a + (int **)&a + +A simple assignment whose left-hand side is a cast works by converting the +right-hand side first to the specified type, then to the type of the inner +left-hand side expression. After this is stored, the value is converter +back to the specified type to become the value of the assignment. Thus, if +`a' has type `char *', the following two expressions are equivalent: + + (int)a = 5 + (int)(a = (char *)5) + +An assignment-with-arithmetic operation such as `+=' applied to a cast +performs the arithmetic using the type resulting from the cast, and then +continues as in the previous case. Therefore, these two expressions are +equivalent: + + (int)a += 5 + (int)(a = (char *) ((int)a + 5)) + + +File: internals, Node: Conditionals, Next: Zero-Length, Prev: Lvalues, Up: Extensions + +Conditional Expressions with Omitted Middle-Operands +==================================================== + +The middle operand in a conditional expression may be omitted. Then if the +first operand is nonzero, its value is the value of the conditional +expression. + +Therefore, the expression + + x ? : y + +has the value of `x' if that is nonzero; otherwise, the value of `y'. + +This example is perfectly equivalent to + + x ? x : y + +In this simple case, the ability to omit the middle operand is not +especially useful. When it becomes useful is when the first operand does, +or may (if it is a macro argument), contain a side effect. Then repeating +the operand in the middle would perform the side effect twice. Omitting +the middle operand uses the value already computed without the undesirable +effects of recomputing it.  -File: internals Node: Comparisons, Prev: Arithmetic, Up: RTL, Next: Bit Fields +File: internals, Node: Zero-Length, Next: Variable-Length, Prev: Conditionals, Up: Extensions -Comparison Operations +Arrays of Length Zero ===================== -Comparison operators test a relation on two operands and are considered to -represent the value 1 if the relation holds, or zero if it does not. The -mode of the comparison is determined by the operands; they must both be -valid for a common machine mode. A comparison with both operands constant -would be invalid as the machine mode could not be deduced from it, but such -a comparison should never exist in rtl due to constant folding. - -Inequality comparisons come in two flavors, signed and unsigned. Thus, -there are distinct expression codes `GT' and `GTU' for signed and -unsigned greater-than. These can produce different results for the same -pair of integer values: for example, 1 is signed greater-than -1 but not -unsigned greater-than, because -1 when regarded as unsigned is actually -0xffffffff which is greater than 1. - -The signed comparisons are also used for floating point values. Floating -point comparisons are distinguished by the machine modes of the operands. - -The comparison operators may be used to compare the condition codes -`(cc0)' against zero, as in `(eq (cc0) (const_int 0))'. -Such a construct actually refers to the result of the preceding -instruction in which the condition codes were set. The above -example stands for 1 if the condition codes were set to say -"zero" or "equal", 0 otherwise. Although the same comparison -operators are used for this as may be used in other contexts -on actual data, no confusion can result since the machine description -would never allow both kinds of uses in the same context. - -`(eq X Y)' - 1 if the values represented by X and Y are equal, - otherwise 0. - -`(ne X Y)' - 1 if the values represented by X and Y are not equal, - otherwise 0. - -`(gt X Y)' - 1 if the X is greater than Y. If they are fixed-point, - the comparison is done in a signed sense. - -`(gtu X Y)' - Like `gt' but does unsigned comparison, on fixed-point numbers only. - -`(lt X Y)' -`(ltu X Y)' - Like `gt' and `gtu' but test for "less than". - -`(ge X Y)' -`(geu X Y)' - Like `gt' and `gtu' but test for "greater than or equal". - -`(le X Y)' -`(leu X Y)' - Like `gt' and `gtu' but test for "less than or equal". - -`(if_then_else COND THEN ELSE)' - This is not a comparison operation but is listed here because it is - always used in conjunction with a comparison operation. To be - precise, COND is a comparison expression. This expression - represents a choice, according to COND, between the value - represented by THEN and the one represented by ELSE. +Zero-length arrays are allowed in GNU C. They are very useful as the last +element of a structure which is really a header for a variable-length object: + + struct line { + int length; + char contents[0]; + }; - On most machines, `if_then_else' expressions are valid only - to express conditional jumps. + { + struct line *thisline + = (struct line *) malloc (sizeof (struct line) + this_length); + thisline->length = thislength; + } + +In standard C, you would have to give `contents' a length of 1, which means +either you waste space or complicate the argument to `malloc'.  -File: internals Node: Bit Fields, Prev: Comparisons, Up: RTL, Next: Conversions +File: internals, Node: Variable-Length, Next: Subscripting, Prev: Zero-Length, Up: Extensions -Bit-fields -========== +Arrays of Variable Length +========================= -Special expression codes exist to represent bit-field instructions. -These types of expressions are lvalues in rtl; they may appear -on the left side of a assignment, indicating insertion of a value -into the specified bit field. - -`(sign_extract:SI LOC SIZE POS)' - This represents a reference to a sign-extended bit-field contained or - starting in LOC (a memory or register reference). The bit field - is SIZE bits wide and starts at bit POS. The compilation - switch `BITS_BIG_ENDIAN' says which end of the memory unit - POS counts from. - - Which machine modes are valid for LOC depends on the machine, - but typically LOC should be a single byte when in memory - or a full word in a register. - -`(zero_extract:SI LOC POS SIZE)' - Like `sign_extract' but refers to an unsigned or zero-extended - bit field. The same sequence of bits are extracted, but they - are filled to an entire word with zeros instead of by sign-extension. - - -File: internals Node: Conversions, Prev: Bit Fields, Up: RTL, Next: RTL Declarations - -Conversions -=========== - -All conversions between machine modes must be represented by -explicit conversion operations. For example, an expression -which the sum of a byte and a full word cannot be written as -`(plus:SI (reg:QI 34) (reg:SI 80))' because the `plus' -operation requires two operands of the same machine mode. -Therefore, the byte-sized operand is enclosed in a conversion -operation, as in - - (plus:SI (sign_extend:SI (reg:QI 34)) (reg:SI 80)) - -The conversion operation is not a mere placeholder, because there -may be more than one way of converting from a given starting mode -to the desired final mode. The conversion operation code says how -to do it. - -`(sign_extend:M X)' - Represents the result of sign-extending the value X - to machine mode M. M must be a fixed-point mode - and X a fixed-point value of a mode narrower than M. - -`(zero_extend:M X)' - Represents the result of zero-extending the value X - to machine mode M. M must be a fixed-point mode - and X a fixed-point value of a mode narrower than M. - -`(float_extend:M X)' - Represents the result of extending the value X - to machine mode M. M must be a floating point mode - and X a floating point value of a mode narrower than M. - -`(truncate:M X)' - Represents the result of truncating the value X - to machine mode M. M must be a fixed-point mode - and X a fixed-point value of a mode wider than M. - -`(float_truncate:M X)' - Represents the result of truncating the value X - to machine mode M. M must be a floating point mode - and X a floating point value of a mode wider than M. - -`(float:M X)' - Represents the result of converting fixed point value X - to floating point mode M. +Variable-length automatic arrays are allowed in GNU C. These arrays are +declared like any other automatic arrays, but with a length that is not a +constant expression. The storage is allocated at that time and deallocated +when the brace-level is exited. For example: + + FILE *concat_fopen (char *s1, char *s2, char *mode) + { + char str[strlen (s1) + strlen (s2) + 1]; + strcpy (str, s1); + strcat (str, s2); + return fopen (str, mode); + } + +You can also define structure types containing variable-length arrays, and +use them even for arguments or function values, as shown here: + + int foo; -`(fix:M X)' - Represents the result of converting floating point value X - to fixed point mode M. How rounding is done is not specified. + struct entry + { + char data[foo]; + }; + struct entry + tester (struct entry arg) + { + struct entry new; + int i; + for (i = 0; i < foo; i++) + new.data[i] = arg.data[i] + 1; + return new; + } + +(Eventually there will be a way to say that the size of the array is +another member of the same structure.) + +The length of an array is computed on entry to the brace-level where the +array is declared and is remembered for the scope of the array in case you +access it with `sizeof'. + +Jumping or breaking out of the scope of the array name will also deallocate +the storage. Jumping into the scope is not allowed; you will get an error +message for it. + +You can use the function `alloca' to get an effect much like +variable-length arrays. The function `alloca' is available in many other C +implementations (but not in all). On the other hand, variable-length +arrays are more elegant. + +There are other differences between these two methods. Space allocated +with `alloca' exists until the containing *function* returns. The space +for a variable-length array is deallocated as soon as the array name's +scope ends. (If you use both variable-length arrays and `alloca' in the +same function, deallocation of a variable-length array will also deallocate +anything more recently allocated with `alloca'.)  -File: internals Node: RTL Declarations, Prev: Conversions, Up: RTL, Next: Side Effects +File: internals, Node: Subscripting, Next: Pointer Arith, Prev: Variable-Length, Up: Extensions -Declarations -============ +Non-Lvalue Arrays May Have Subscripts +===================================== -Declaration expression codes do not represent arithmetic operations -but rather state assertions about their operands. +Subscripting is allowed on arrays that are not lvalues, even though the +unary `&' operator is not. For example, this is valid in GNU C though not +valid in other C dialects: -`(volatile:M X)' - Represents the same value X does, but makes the assertion - that it should be treated as a volatile value. This forbids - coalescing multiple accesses or deleting them even if it would - appear to have no effect on the program. X must be a `mem' - expression with mode M. - - The first thing the reload pass does to an insn is to remove all - `volatile' expressions from it; each one is replaced by its - operand. + struct foo {int a[4];}; - Recognizers will never recognize anything with `volatile' in it. - This automatically prevents some optimizations on such things - (such as instruction combination). After the reload pass removes - all volatility information, the insns can be recognized. + struct foo f(); - Cse removes `volatile' from destinations of `set''s, because - no optimizations reorder such `set's. This is not required for - correct code and is done to permit some optimization on the value to - be stored. - -`(unchanging:M X)' - Represents the same value X does, but makes the assertion - that its value is effectively constant during the execution - of the current function. This permits references to X - to be moved freely within the function. X must be a `reg' - expression with mode M. - -`(strict_low_part (subreg:M (reg:N R) 0))' - This expression code is used in only one context: operand 0 of a - `set' expression. In addition, the operand of this expression - must be a `subreg' expression. - - The presence of `strict_low_part' says that the part of the - register which is meaningful in mode N but is not part of - mode M is not to be altered. Normally, an assignment to such - a subreg is allowed to have undefined effects on the rest of the - register when M is less than a word. + bar (int index) + { + return f().a[index]; + } + + +File: internals, Node: Pointer Arith, Next: Initializers, Prev: Subscripting, Up: Extensions + +Arithmetic on `void'-Pointers and Function Pointers +=================================================== + +In GNU C, addition and subtraction operations are supported on pointers to +`void' and on pointers to functions. This is done by treating the size of +a `void' or of a function as 1. + +A consequence of this is that `sizeof' is also allowed on `void' and on +function types, and returns 1. + + +File: internals, Node: Initializers, Next: Constructors, Prev: Pointer Arith, Up: Extensions + +Non-Constant Initializers +========================= + +The elements of an aggregate initializer are not required to be constant +expressions in GNU C. Here is an example of an initializer with run-time +varying elements: + + foo (float f, float g) + { + float beat_freqs[2] = { f-g, f+g }; + ... + }  -File: internals Node: Side Effects, Prev: RTL Declarations, Up: RTL, Next: Incdec +File: internals, Node: Constructors, Next: Dollar Signs, Prev: Initializers, Up: Extensions -Side Effect Expressions +Constructor Expressions ======================= -The expression codes described so far represent values, not actions. -But machine instructions never produce values; they are meaningful -only for their side effects on the state of the machine. Special -expression codes are used to represent side effects. - -The body of an instruction is always one of these side effect codes; -the codes described above, which represent values, appear only as -the operands of these. - -`(set LVAL X)' - Represents the action of storing the value of X into the place - represented by LVAL. LVAL must be an expression - representing a place that can be stored in: `reg' (or - `subreg' or `strict_low_part'), `mem', `pc' or - `cc0'. - - If LVAL is a `reg', `subreg' or `mem', it has a - machine mode; then X must be valid for that mode. - - If LVAL is a `reg' whose machine mode is less than the full - width of the register, then it means that the part of the register - specified by the machine mode is given the specified value and the - rest of the register receives an undefined value. Likewise, if - LVAL is a `subreg' whose machine mode is narrower than - `SImode', the rest of the register can be changed in an undefined way. - - If LVAL is a `strict_low_part' of a `subreg', then the - part of the register specified by the machine mode of the - `subreg' is given the value X and the rest of the register - is not changed. - - If LVAL is `(cc0)', it has no machine mode, and X may - have any mode. This represents a "test" or "compare" instruction. - - If LVAL is `(pc)', we have a jump instruction, and the - possibilities for X are very limited. It may be a - `label_ref' expression (unconditional jump). It may be an - `if_then_else' (conditional jump), in which case either the - second or the third operand must be `(pc)' (for the case which - does not jump) and the other of the two must be a `label_ref' - (for the case which does jump). X may also be a `mem' or - `(plus:SI (pc) Y)', where Y may be a `reg' or a - `mem'; these unusual patterns are used to represent jumps through - branch tables. - -`(return)' - Represents a return from the current function, on machines where - this can be done with one instruction, such as Vaxen. On machines - where a multi-instruction "epilogue" must be executed in order - to return from the function, returning is done by jumping to a - label which precedes the epilogue, and the `return' expression - code is never used. - -`(call FUNCTION NARGS)' - Represents a function call. FUNCTION is a `mem' expression - whose address is the address of the function to be called. NARGS - is an expression representing the number of words of argument. - - Each machine has a standard machine mode which FUNCTION must - have. The machine descripion defines macro `FUNCTION_MODE' to - expand into the requisite mode name. The purpose of this mode is to - specify what kind of addressing is allowed, on machines where the - allowed kinds of addressing depend on the machine mode being - addressed. - -`(clobber X)' - Represents the storing or possible storing of an unpredictable, - undescribed value into X, which must be a `reg' or - `mem' expression. - - One place this is used is in string instructions that store standard - values into particular hard registers. It may not be worth the - trouble to describe the values that are stored, but it is essential - to inform the compiler that the registers will be altered, lest it - attempt to keep data in them across the string instruction. - - X may also be null---a null C pointer, no expression at all. - Such a `(clobber (null))' expression means that all memory - locations must be presumed clobbered. - - Note that the machine description classifies certain hard registers as - "call-clobbered". All function call instructions are assumed by - default to clobber these registers, so there is no need to use - `clobber' expressions to indicate this fact. Also, each function - call is assumed to have the potential to alter any memory location. - -`(use X)' - Represents the use of the value of X. It indicates that - the value in X at this point in the program is needed, - even though it may not be apparent whythis is so. Therefore, the - compiler will not attempt to delete instructions whose only - effect is to store a value in X. X must be a `reg' - expression. - -`(parallel [X0 X1 ...])' - Represents several side effects performed in parallel. The square - brackets stand for a vector; the operand of `parallel' is a - vector of expressions. X0, X1 and so on are individual - side effects---expressions of code `set', `call', - `return', `clobber' or `use'. - - "In parallel" means that first all the values used in - the individual side-effects are computed, and second all the actual - side-effects are performed. For example, - - (parallel [(set (reg:SI 1) (mem:SI (reg:SI 1))) - (set (mem:SI (reg:SI 1)) (reg:SI 1))]) - - says unambiguously that the values of hard register 1 and the memory - location addressed by it are interchanged. In both places where - `(reg:SI 1)' appears as a memory address it refers to the value - in register 1 before the execution of the instruction. - -Three expression codes appear in place of a side effect, as the body -of an insn, though strictly speaking they do not describe side effects -as such: +GNU C supports constructor expressions. A constructor looks like a cast +containing an initializer. Its value is an object of the type specified in +the cast, containing the elements specified in the initializer. The type +must be a structure, union or array type. -`(asm_input S)' - Represents literal assembler code as described by the string S. - -`(addr_vec:M [LR0 LR1 ...])' - Represents a table of jump addresses. LR0 etc. are - `label_ref' expressions. The mode M specifies how much - space is given to each address; normally M would be - `Pmode'. - -`(addr_diff_vec:M BASE [LR0 LR1 ...])' - Represents a table of jump addresses expressed as offsets from - BASE. LR0 etc. are `label_ref' expressions and so is - BASE. The mode M specifies how much space is given to - each address-difference. - - -File: internals Node: Incdec, Prev: Side Effects, Up: RTL, Next: Insns - -Embedded Side-Effects on Addresses -================================== - -Four special side-effect expression codes appear as memory addresses. - -`(pre_dec:M X)' - Represents the side effect of decrementing X by a standard - amount and represents also the value that X has after being - decremented. X must be a `reg' or `mem', but most - machines allow only a `reg'. M must be the machine mode - for pointers on the machine in use. The amount X is decrement - by is the length in bytes of the machine mode of the containing memory - reference of which this expression serves as the address. Here is an - example of its use: - - (mem:DF (pre_dec:SI (reg:SI 39))) - - This says to decrement pseudo register 39 by the length of a `DFmode' - value and use the result to address a `DFmode' value. - -`(pre_inc:M X)' - Similar, but specifies incrementing X instead of decrementing it. - -`(post_dec:M X)' - Represents the same side effect as `pre_decrement' but a different - value. The value represented here is the value X has before - being decremented. - -`(post_inc:M X)' - Similar, but specifies incrementing X instead of decrementing it. +Assume that `struct foo' and `structure' are declared as shown: -These embedded side effect expressions must be used with care. Instruction -patterns may not use them. Until the `flow' pass of the compiler, -they may occur only to represent pushes onto the stack. The `flow' -pass finds cases where registers are incremented or decremented in one -instruction and used as an address shortly before or after; these cases are -then transformed to use pre- or post-increment or -decrement. - -Explicit popping of the stack could be represented with these embedded -side effect operators, but that would not be safe; the instruction -combination pass could move the popping past pushes, thus changing -the meaning of the code. - -An instruction that can be represented with an embedded side effect -could also be represented using `parallel' containing an additional -`set' to describe how the address register is altered. This is not -done because machines that allow these operations at all typically -allow them wherever a memory address is called for. Describing them as -additional parallel stores would require doubling the number of entries -in the machine description. - - -File: internals Node: Insns, Prev: Incdec, Up: RTL, Next: Sharing - -Insns -===== - -The RTL representation of the code for a function is a doubly-linked -chain of objects called "insns". Insns are expressions with -special codes that are used for no other purpose. Some insns are -actual instructions; others represent dispatch tables for `switch' -statements; others represent labels to jump to or various sorts of -declaratory information. - -In addition to its own specific data, each insn must have a unique id number -that distinguishes it from all other insns in the current function, and -chain pointers to the preceding and following insns. These three fields -occupy the same position in every insn, independent of the expression code -of the insn. They could be accessed with `XEXP' and `XINT', -but instead three special macros are always used: + struct foo {int a; char b[2];} structure; -`INSN_UID (I)' - Accesses the unique id of insn I. - -`PREV_INSN (I)' - Accesses the chain pointer to the insn preceding I. - If I is the first insn, this is a null pointer. - -`NEXT_INSN (I)' - Accesses the chain pointer to the insn following I. - If I is the last insn, this is a null pointer. +Here is an example of constructing a `struct foo' with a constructor: -The `NEXT_INSN' and `PREV_INSN' pointers must always -correspond: if I is not the first insn, + structure = ((struct foo) {x + y, 'a', 0}); - NEXT_INSN (PREV_INSN (INSN)) == INSN +This is equivalent to writing the following: -is always true. + { + struct foo temp = {x + y, 'a', 0}; + structure = temp; + } -Every insn has one of the following six expression codes: +You can also construct an array. If all the elements of the constructor +are (made up of) simple constant expressions, suitable for use in +initializers, then the constructor is an lvalue and can be coerced to a +pointer to its first element, as shown here: -`insn' - The expression code `insn' is used for instructions that do not jump - and do not do function calls. Insns with code `insn' have four - additional fields beyond the three mandatory ones listed above. - These four are described in a table below. - -`jump_insn' - The expression code `jump_insn' is used for instructions that may jump - (or, more generally, may contain `label_ref' expressions). - `jump_insn' insns have the same extra fields as `insn' insns, - accessed in the same way. - -`call_insn' - The expression code `call_insn' is used for instructions that may do - function calls. It is important to distinguish these instructions because - they imply that certain registers and memory locations may be altered - unpredictably. - - `call_insn' insns have the same extra fields as `insn' insns, - accessed in the same way. - -`code_label' - A `code_label' insn represents a label that a jump insn can jump to. - It contains one special field of data in addition to the three standard ones. - It is used to hold the "label number", a number that identifies this - label uniquely among all the labels in the compilation (not just in the - current function). Ultimately, the label is represented in the assembler - output as an assembler label `LN' where N is the label number. - -`barrier' - Barriers are placed in the instruction stream after unconditional - jump instructions to indicate that the jumps are unconditional. - They contain no information beyond the three standard fields. - -`note' - `note' insns are used to represent additional debugging and - declaratory information. They contain two nonstandard fields, an - integer which is accessed with the macro `NOTE_LINE_NUMBER' and a - string accessed with `NOTE_SOURCE_FILE'. - - If `NOTE_LINE_NUMBER' is positive, the note represents the - position of a source line and `NOTE_SOURCE_FILE' is the source file name - that the line came from. These notes control generation of line - number data in the assembler output. - - Otherwise, `NOTE_LINE_NUMBER' is not really a line number but a - code with one of the following values (and `NOTE_SOURCE_FILE' - must contain a null pointer): - - `NOTE_INSN_DELETED' - Such a note is completely ignorable. Some passes of the compiler - delete insns by altering them into notes of this kind. - - `NOTE_INSN_BLOCK_BEG' - `NOTE_INSN_BLOCK_END' - These types of notes indicate the position of the beginning and end - of a level of scoping of variable names. They control the output - of debugging information. - - `NOTE_INSN_LOOP_BEG' - `NOTE_INSN_LOOP_END' - These types of notes indicate the position of the beginning and end - of a `while' or `for' loop. They enable the loop optimizer - to find loops quickly. + char **foo = (char *[]) { "x", "y", "z" }; -Here is a table of the extra fields of `insn', `jump_insn' -and `call_insn' insns: +Array constructors whose elements are not simple constants are not very +useful, because the constructor is not an lvalue. There are only two valid +ways to use it: to subscript it, or initialize an array variable with it. +The former is probably slower than a `switch' statement, while the latter +does the same thing an ordinary C initializer would do. -`PATTERN (I)' - An expression for the side effect performed by this insn. - -`REG_NOTES (I)' - A list (chain of `expr_list' expressions) giving information - about the usage of registers in this insn. This list is set up by the - `flow' pass; it is a null pointer until then. - -`LOG_LINKS (I)' - A list (chain of `insn_list' expressions) of previous "related" - insns: insns which store into registers values that are used for the - first time in this insn. (An additional constraint is that neither a - jump nor a label may come between the related insns). This list is - set up by the `flow' pass; it is a null pointer until then. - -`INSN_CODE (I)' - An integer that says which pattern in the machine description matches - this insn, or -1 if the matching has not yet been attempted. - - Such matching is never attempted and this field is not used on an insn - whose pattern consists of a single `use', `clobber', - `asm', `addr_vec' or `addr_diff_vec' expression. - -The `LOG_LINKS' field of an insn is a chain of `insn_list' -expressions. Each of these has two operands: the first is an insn, -and the second is another `insn_list' expression (the next one in -the chain). The last `insn_list' in the chain has a null pointer -as second operand. The significant thing about the chain is which -insns apepar in it (as first operands of `insn_list' -expressions). Their order is not significant. - -The `REG_NOTES' field of an insn is a similar chain but of -`expr_list' expressions instead of `insn_list'. The first -operand is a `reg' rtx. Its presence in the list can have three -possible meanings, distinguished by a value that is stored in the -machine-mode field of the `expr_list' because that is a -conveniently available space, but that is not really a machine mode. -These values belong to the C type `enum reg_note' and there are -three of them: - -`REG_DEAD' - The `reg' listed dies in this insn; that is to say, altering - the value immediately after this insn would not affect the future - behavior of the program. - -`REG_INC' - The `reg' listed is incremented (or decremented; at this level - there is no distinction) by an embedded side effect inside this insn. - -`REG_CONST' - The `reg' listed has a value that could safely be replaced - everywhere by the value that this insn copies into it. ("Safety" - here refers to the data flow of the program; such replacement may - require reloading into registers for some of the insns in which - the `reg' is replaced.) - -`REG_WAS_0' - The `reg' listed contained zero before this insn. You can rely - on this note if it is present; its absence implies nothing. - -(The only difference between the expression codes `insn_list' and -`expr_list' is that the first operand of an `insn_list' is -assumed to be an insn and is printed in debugging dumps as the insn's -unique id; the first operand of an `expr_list' is printed in the -ordinary way as an expression.) - - -File: internals Node: Sharing, Prev: Insns, Up: RTL - -Structure Sharing Assumptions -============================= - -The compiler assumes that certain kinds of RTL expressions are unique; -there do not exist two distinct objects representing the same value. -In other cases, it makes an opposite assumption: that no RTL expression -object of a certain kind appears in more than one place in the -containing structure. - -These assumptions refer to a single function; except for the RTL -objects that describe global variables and external functions, -no RTL objects are common to two functions. + output = ((int[]) { 2, x, 28 }) [input]; - * Each pseudo-register has only a single `reg' object to represent it, - and therefore only a single machine mode. - - * For any symbolic label, there is only one `symbol_ref' object - referring to it. - - * There is only one `const_int' expression with value zero, - and only one with value one. - - * There is only one `pc' expression. - - * There is only one `cc0' expression. - - * There is only one `const_double' expression with mode - `SFmode' and value zero, and only one with mode `DFmode' and - value zero. - - * No `label_ref' appears in more than one place in the RTL structure; - in other words, it is safe to do a tree-walk of all the insns in the function - and assume that each time a `label_ref' is seen it is distinct from all - other `label_refs' seen. - - * Aside from the cases listed above, the only kind of expression - object that may appear in more than one place is the `mem' - object that describes a stack slot or a static variable. + +File: internals, Node: Dollar Signs, Next: Alignment, Prev: Constructors, Up: Extensions + +Dollar Signs in Identifier Names +================================ + +In GNU C, you may use dollar signs in identifier names. This is because +many traditional C implementations allow such identifiers.  -File: internals Node: Machine Desc, Prev: RTL, Up: Top, Next: Machine Macros +File: internals, Node: Alignment, Next: Inline, Prev: Dollar Signs, Up: Extensions -Machine Descriptions -******************** +Inquiring about the Alignment of a Type or Variable +=================================================== -A machine description has two parts: a file of instruction patterns -(`.md' file) and a C header file of macro definitions. +The keyword `__alignof' allows you to inquire about how an object is +aligned, or the minimum alignment usually required by a type. Its syntax +is just like `sizeof'. -The `.md' file for a target machine contains a pattern for each -instruction that the target machine supports (or at least each instruction -that is worth telling the compiler about). It may also contain comments. -A semicolon causes the rest of the line to be a comment, unless the semicolon -is inside a quoted string. +For example, if the target machine requires a `double' value to be aligned +on an 8-byte boundary, then `__alignof (double)' is 8. This is true on +many RISC machines. On more traditional machine designs, `__alignof +(double)' is 4 or even 2. -See the next chapter for information on the C header file. +Some machines never actually require alignment; they allow reference to any +data type even at an odd addresses. For these machines, `__alignof' +reports the *recommended* alignment of a type. -* Menu: +When the operand of `__alignof' is an lvalue rather than a type, the value +is the largest alignment that the lvalue is known to have. It may have +this alignment as a result of its data type, or because it is part of a +structure and inherits alignment from that structure. For example, after +this declaration: -* Patterns:: How to write instruction patterns. -* Example:: Example of an instruction pattern. -* Constraints:: When not all operands are general operands. -* Standard Names:: Names mark patterns to use for code generation. -* Dependent Patterns:: Having one pattern may make you need another. - - -File: internals Node: Patterns, Prev: Machine Desc, Up: Machine Desc, Next: Example - -Instruction Patterns -==================== - -Each instruction pattern contains an incomplete RTL expression, with pieces -to be filled in later, operand constraints that restrict how the pieces can -be filled in, and an output pattern or C code to generate the assembler -output, all wrapped up in a `define_insn' expression. - -Sometimes an insn can match more than one instruction pattern. Then the -pattern that appears first in the machine description is the one used. -Therefore, more specific patterns should usually go first in the -description. - -The `define_insn' expression contains four operands: - - 1. An optional name. The presence of a name indicate that this instruction - pattern can perform a certain standard job for the RTL-generation - pass of the compiler. This pass knows certain names and will use - the instruction patterns with those names, if the names are defined - in the machine description. - - The absence of a name is indicated by writing an empty string - where the name should go. Nameless instruction patterns are never - used for generating RTL code, but they may permit several simpler insns - to be combined later on. - - Names that are not thus known and used in RTL-generation have no - effect; they are equivalent to no name at all. - - 2. The recognition template. This is a vector of incomplete RTL - expressions which show what the instruction should look like. It is - incomplete because it may contain `match_operand' and - `match_dup' expressions that stand for operands of the - instruction. - - If the vector has only one element, that element is what the - instruction should look like. If the vector has multiple elements, - then the instruction looks like a `parallel' expression - containing that many elements as described. - - 3. A condition. This is a string which contains a C expression that is - the final test to decide whether an insn body matches this pattern. - - For a named pattern, the condition (if present) may not depend on - the data in the insn being matched, but only the target-machine-type - flags. The compiler needs to test these conditions during - initialization in order to learn exactly which named instructions are - available in a particular run. - - For nameless patterns, the condition is applied only when matching an - individual insn, and only after the insn has matched the pattern's - recognition template. The insn's operands may be found in the vector - `operands'. - - 4. A string that says how to output matching insns as assembler code. In - the simpler case, the string is an output template, much like a - `printf' control string. `%' in the string specifies where - to insert the operands of the instruction; the `%' is followed by - a single-digit operand number. - - `%cDIGIT' can be used to subtitute an operand that is a - constant value without the syntax that normally indicates an immediate - operand. - - `%aDIGIT' can be used to substitute an operand as if it - were a memory reference, with the actual operand treated as the address. - This may be useful when outputting a "load address" instruction, - because often the assembler syntax for such an instruction requires - you to write the operand as if it were a memory reference. - - The template may generate multiple assembler instructions. - Write the text for the instructions, with `\;' between them. - - If the output control string starts with a `*', then it is not an - output template but rather a piece of C program that should compute a - template. It should execute a `return' statement to return the - template-string you want. Most such templates use C string literals, - which require doublequote characters to delimit them. To include - these doublequote characters in the string, prefix each one with - `\'. - - The operands may be found in the array `operands', whose C - data type is `rtx []'. - - It is possible to output an assembler instruction and then go on to - output or compute more of them, using the subroutine - `output_asm_insn'. This receives two arguments: a - template-string and a vector of operands. The vector may be - `operands', or it may be another array of `rtx' that you - declare locally and initialize yourself. - -The recognition template is used also, for named patterns, for -constructing insns. Construction involves substituting specified -operands into a copy of the template. Matching involves determining -the values that serve as the operands in the insn being matched. Both -of these activities are controlled by two special expression types -that direct matching and substitution of the operands. - -`(match_operand:M N TESTFN CONSTRAINT)' - This expression is a placeholder for operand number N of - the insn. When constructing an insn, operand number N - will be substituted at this point. When matching an insn, whatever - appears at this position in the insn will be taken as operand - number N; but it must satisfy TESTFN or this instruction - pattern will not match at all. - - Operand numbers must be chosen consecutively counting from zero in - each instruction pattern. There may be only one `match_operand' - expression in the pattern for each expression number, and they must - appear in order of increasing expression number. - - TESTFN is a string that is the name of a C function that accepts - two arguments, a machine mode and an expression. During matching, - the function will be called with M as the mode argument - and the putative operand as the other argument. If it returns zero, - this instruction pattern fails to match. TESTFN may be - an empty string; then it means no test is to be done on the operand. - - Most often, TESTFN is `"general_operand"'. It checks - that the putative operand is either a constant, a register or a - memory reference, and that it is valid for mode M. - - CONSTRAINT is explained later. - -`(match_dup N)' - This expression is also a placeholder for operand number N. - It is used when the operand needs to appear more than once in the - insn. - - In construction, `match_dup' behaves exactly like - MATCH_OPERAND: the operand is substituted into the insn being - constructed. But in matching, `match_dup' behaves differently. - It assumes that operand number N has already been determined by - a `match_operand' apparing earlier in the recognition template, - and it matches only an identical-looking expression. - -`(address (match_operand:M N "address_operand" ""))' - This complex of expressions is a placeholder for an operand number - N in a "load address" instruction: an operand which specifies - a memory location in the usual way, but for which the actual operand - value used is the address of the location, not the contents of the - location. - - `address' expressions never appear in RTL code, only in machine - descriptions. And they are used only in machine descriptions that do - not use the operand constraint feature. When operand constraints are - in use, the letter `p' in the constraint serves this purpose. - - M is the machine mode of the *memory location being - addressed*, not the machine mode of the address itself. That mode is - always the same on a given target machine (it is `Pmode', which - normally is `SImode'), so there is no point in mentioning it; - thus, no machine mode is written in the `address' expression. If - some day support is added for machines in which addresses of different - kinds of objects appear differently or are used differently (such as - the PDP-10), different formats would perhaps need different machine - modes and these modes might be written in the `address' - expression. - - -File: internals Node: Example, Prev: Patterns, Up: Machine Desc, Next: Constraints - -Example of `define_insn' -======================== - -Here is an actual example of an instruction pattern, for the 68000/68020. - - (define_insn "tstsi" - [(set (cc0) - (match_operand:SI 0 "general_operand" "rm"))] - "" - "* - { if (TARGET_68020 || ! ADDRESS_REG_P (operands[0])) - return \"tstl %0\"; - return \"cmpl #0,%0\"; }") - -This is an instruction that sets the condition codes based on the value of -a general operand. It has no condition, so any insn whose RTL description -has the form shown may be handled according to this pattern. The name -`tstsi' means "test a `SImode' value" and tells the RTL generation -pass that, when it is necessary to test such a value, an insn to do so -can be constructed using this pattern. - -The output control string is a piece of C code which chooses which -output template to return based on the kind of operand and the specific -type of CPU for which code is being generated. + struct foo { int x; char y; } foo1; -`"rm"' is an operand constraint. Its meaning is explained below. +the value of `__alignof (foo1.y)' is probably 2 or 4, the same as +`__alignof (int)', even though the data type of `foo1.y' does not itself +demand any alignment.  -File: internals Node: Constraints, Prev: Example, Up: Machine Desc, Next: Standard Names +File: internals, Node: Inline, Next: Extended Asm, Prev: Alignment, Up: Extensions -Operand Constraints -=================== +An Inline Function is As Fast As a Macro +======================================== -Each `match_operand' in an instruction pattern can specify a -constraint for the type of operands allowed. Constraints can say whether -an operand may be in a register, and which kinds of register; whether the -operand can be a memory reference, and which kinds of address; whether the -operand may be an immediate constant, and which possible values it may -have. Constraints can also require two operands to match. +By declaring a function `inline', you can direct GNU CC to integrate that +function's code into the code for its callers. This makes execution faster +by eliminating the function-call overhead; in addition, if any of the +actual argument values are constant, their known values may permit +simplifications at compile time so that not all of the inline function's +code needs to be included. -* Menu: +To declare a function inline, use the `inline' keyword in its declaration, +like this: -* Simple Constraints:: Basic use of constraints. -* Multi-alternative:: When an insn has two alternative constraint-patterns. -* Class Preferences:: Constraints guide which hard register to put things in. -* Modifiers:: More precise control over effects of constraints. -* No Constraints:: Describing a clean machine without constraints. + inline int + inc (int *a) + { + (*a)++; + } + +You can also make all ``simple enough'' functions inline with the option +`-finline-functions'. Note that certain usages in a function definition +can make it unsuitable for inline substitution. + +When a function is both inline and `static', if all calls to the function +are integrated into the caller, then the function's own assembler code is +never referenced. In this case, GNU CC does not actually output assembler +code for the function, unless you specify the option +`-fkeep-inline-functions'. Some calls cannot be integrated for various +reasons (in particular, calls that precede the function's definition cannot +be integrated, and neither can recursive calls within the definition). If +there is a nonintegrated call, then the function is compiled to assembler +code as usual. + +When an inline function is not `static', then the compiler must assume that +there may be calls from other source files; since a global symbol can be +defined only once in any program, the function must not be defined in the +other source files, so the calls therein cannot be integrated. Therefore, +a non-`static' inline function is always compiled on its own in the usual +fashion.  -File: internals Node: Simple Constraints, Prev: Constraints, Up: Constraints, Next: Multi-Alternative +File: internals, Node: Extended Asm, Next: Asm Labels, Prev: Inline, Up: Extensions -Simple Constraints ------------------- +Assembler Instructions with C Expression Operands +================================================= -The simplest kind of constraint is a string full of letters, each of -which describes one kind of operand that is permitted. Here are -the letters that are allowed: +In an assembler instruction using `asm', you can now specify the operands +of the instruction using C expressions. This means no more guessing which +registers or memory locations will contain the data you want to use. -`m' - A memory operand is allowed, with any kind of address that the machine - supports in general. - -`o' - A memory operand is allowed, but only if the address is "offsetable". - This means that adding a small integer (actually, the width in bytes of the - operand, as determined by its machine mode) may be added to the address - and the result is also a valid memory address. For example, an address - which is constant is offsetable; so is an address that is the sum of - a register and a constant (as long as a slightly larger constant is also - within the range of address-offsets supported by the machine); but an - autoincrement or autodecrement address is not offsetable. More complicated - indirect/indexed addresses may or may not be offsetable depending on the - other addressing modes that the machine supports. - -`<' - A memory operand with autodecrement addressing (either predecrement or - postdecrement) is allowed. - -`>' - A memory operand with autoincrement addressing (either preincrement or - postincrement) is allowed. - -`r' - A register operand is allowed provided that it is in a general register. - -`d' -`a' -`f' -`...' - Other letters can be defined in machine-dependent fashion to stand for - particular classes of registers. `d', `a' and `f' are - defined on the 68000/68020 to stand for data, address and floating point - registers. - -`i' - An immediate integer operand (one with constant value) is allowed. - -`I' -`J' -`K' -`...' - Other letters in the range `I' through `M' may be defined in a - machine-dependent fashion to permit immediate integer operands with - explicit integer values in specified ranges. For example, on the 68000, - `I' is defined to stand for the range of values 1 to 8. This is the - range permitted as a shift count in the shift instructions. - -`F' - An immediate floating operand (expression code `const_double') is - allowed. - -`G' -`H' - `G' and `H' may be defined in a machine-dependent fashion to - permit immediate floating operands in particular ranges of values. - -`s' - An immediate integer operand whose value is not an explicit integer is - allowed. This might appear strange; if an insn allows a constant operand - with a value not known at compile time, it certainly must allow any known - value. So why use `s' instead of `i'? Sometimes it allows - better code to be generated. For example, on the 68000 in a fullword - instruction it is possible to use an immediate operand; but if the - immediate value is between -32 and 31, better code results from loading the - value into a register and using the register. This is because the load - into the register can be done with a `moveq' instruction. We arrange - for this to happen by defining the letter `K' to mean "any integer - outside the range -32 to 31", and then specifying `Ks' in the operand - constraints. - -`g' - Any register, memory or immediate integer operand is allowed, except for - registers that are not general registers. - -`N, a digit' - An operand identical to operand number N is allowed. - If a digit is used together with letters, the digit should come last. - -`p' - An operand that is a valid memory address is allowed. This is - for "load address" and "push address" instructions. - - If `p' is used in the constraint, the test-function in the - `match_operand' must be `address_operand'. +You must specify an assembler instruction template much like what appears +in a machine description, plus an operand constraint string for each operand. -In order to have valid assembler code, each operand must satisfy -its constraint. But a failure to do so does not prevent the pattern -from applying to an insn. Instead, it directs the compiler to modify -the code such that the constraint will be satisfied. Usually this is -done by copying an operand into a register. - -Contrast, therefore, the two instruction patterns that follow: - - (define_insn "" - [(set (match_operand:SI 0 "general_operand" "r") - (plus:SI (match_dup 0) - (match_operand:SI 1 "general_operand" "r")))] - "" - "...") - -which has two operands, one of which must appear in two places, and - - (define_insn "" - [(set (match_operand:SI 0 "general_operand" "r") - (plus:SI (match_operand:SI 1 "general_operand" "0") - (match_operand:SI 2 "general_operand" "r")))] - "" - "...") - -which has three operands, two of which are required by a constraint to be -identical. If we are considering an insn of the form - - (insn N PREV NEXT - (set (reg:SI 3) - (plus:SI (reg:SI 6) (reg:SI 109))) - ...) - -the first pattern would not apply at all, because this insn does not -contain two identical subexpressions in the right place. The pattern would -say, "That does not look like an add instruction; try other patterns." -The second pattern would say, "Yes, that's an add instruction, but there -is something wrong with it." It would direct the reload pass of the -compiler to generate additional insns to make the constraint true. The -results might look like this: - - (insn N2 PREV N - (set (reg:SI 3) (reg:SI 6)) - ...) - - (insn N N2 NEXT - (set (reg:SI 3) - (plus:SI (reg:SI 3) (reg:SI 109))) - ...) - -Because insns that don't fit the constraints are fixed up by loading -operands into registers, every instruction pattern's constraints must -permit the case where all the operands are in registers. It need not -permit all classes of registers; the compiler knows how to copy registers -into other registers of the proper class in order to make an instruction -valid. But if no registers are permitted, the compiler will be stymied: it -does not know how to save a register in memory in order to make an -instruction valid. Instruction patterns that reject registers can be -made valid by attaching a condition-expression that refuses to match -an insn at all if the crucial operand is a register. - - -File: internals Node: Multi-Alternative, Prev: Simple Constraints, Up: Constraints, Next: Class Preferences - -Multiple Alternative Constraints --------------------------------- - -Sometimes a single instruction has multiple alternative sets of possible -operands. For example, on the 68000, a logical-or instruction can combine -register or an immediate value into memory, or it can combine any kind of -operand into a register; but it cannot combine one memory location into -another. - -These constraints are represented as multiple alternatives. An alternative -can be described by a series of letters for each operand. The overall -constraint for an operand is made from the letters for this operand -from the first alternative, a comma, the letters for this operand from -the second alternative, a comma, and so on until the last alternative. -Here is how it is done for fullword logical-or on the 68000: - - (define_insn "iorsi3" - [(set (match_operand:SI 0 "general_operand" "=%m,d") - (ior:SI (match_operand:SI 1 "general_operand" "0,0") - (match_operand:SI 2 "general_operand" "dKs,dmKs")))] - ...) - -The first alternative has `m' (memory) for operand 0, `0' for -operand 1 (meaning it must match operand 0), and `dKs' for operand 2. -The second alternative has `d' (data register) for operand 0, `0' -for operand 1, and `dmKs' for operand 2. The `=' and `%' in -the constraint for operand 0 are not part of any alternative; their meaning -is explained in the next section. - -If all the operands fit any one alternative, the instruction is valid. -Otherwise, for each alternative, the compiler counts how many instructions -must be added to copy the operands so that that alternative applies. -The alternative requiring the least copying is chosen. If two alternatives -need the same amount of copying, the one that comes first is chosen. -These choices can be altered with the `?' and `!' characters: - -`?' - Disparage slightly the alternative that the `?' appears in, - as a choice when no alternative applies exactly. The compiler regards - this alternative as one unit more costly for each `?' that appears - in it. - -`!' - Disparage severely the alternative that the `!' appears in. - When operands must be copied into registers, the compiler will - never choose this alternative as the one to strive for. +For example, here is how to use the 68881's `fsinx' instruction: - -File: internals Node: Class Preferences, Prev: Multi-Alternative, Up: Constraints, Next: Modifiers + asm ("fsinx %1,%0" : "=f" (result) : "f" (angle)); + +Here `angle' is the C expression for the input operand while `result' is +that of the output operand. Each has `"f"' as its operand constraint, +saying that a floating-point register is required. The constraints use the +same language used in the machine description (*Note Constraints::.). + +Each operand is described by an operand-constraint string followed by the C +expression in parentheses. A colon separates the assembler template from +the first output operand, and another separates the last output operand +from the first input, if any. Commas separate output operands and separate +inputs. The number of operands is limited to the maximum number of +operands in any instruction pattern in the machine description. + +Output operand expressions must be lvalues, and there must be at least one +of them. The compiler can check this. The input operands need not be +lvalues, and there need not be any. The compiler cannot check whether the +operands have data types that are reasonable for the instruction being +executed. + +The output operands must be write-only; GNU CC will assume that the values +in these operands before the instruction are dead and need not be +generated. For an operand that is read-write, you must logically split its +function into two separate operands, one input operand and one write-only +output operand. The connection between them is expressed by constraints +which say they need to be in the same location when the instruction +executes. You can use the same C expression for both operands, or +different expressions. For example, here we write the (fictitious) +`combine' instruction with `bar' as its read-only source operand and `foo' +as its read-write destination: + + asm ("combine %2,%0" : "=r" (foo) : "0" (foo), "g" (bar)); + +The constraint `"0"' for operand 1 says that it must occupy the same +location as operand 0. Therefore it is not necessary to substitute operand +1 into the assembler code output. -Register Class Preferences --------------------------- +Usually the most convenient way to use these `asm' instructions is to +encapsulate them in macros that look like functions. For example, -The operand constraints have another function: they enable the compiler -to decide which kind of hardware register a pseudo register is best -allocated to. The compiler examines the constraints that apply to the -insns that use the pseudo register, looking for the machine-dependent -letters such as `d' and `a' that specify classes of registers. -The pseudo register is put in whichever class gets the most "votes". -The constraint letters `g' and `r' also vote: they vote in -favor of a general register. The machine description says which registers -are considered general. + #define sin(x) \ + ({ double __value, __arg = (x); \ + asm ("fsinx %1,%0": "=f" (__value): "f" (__arg)); \ + __value; }) -Of course, on some machines all registers are equivalent, and no register -classes are defined. Then none of this complexity is relevant. +Here the variable `__arg' is used to make sure that the instruction +operates on a proper `double' value, and to accept only those arguments `x' +which can convert automatically to a `double'. + +Another way to make sure the instruction operates on the correct data type +is to use a cast in the `asm'. This is different from using a variable +`__arg' in that it converts more different types. For example, if the +desired type were `int', casting the argument to `int' would accept a +pointer with no complaint, while assigning the argument to an `int' +variable named `__arg' would warn about using a pointer unless the caller +explicitly casts it. + +GNU CC assumes for optimization purposes that these instructions have no +side effects except to change the output operands. This does not mean that +instructions with a side effect cannot be used, but you must be careful, +because the compiler may eliminate them if the output operands aren't used, +or move them out of loops, or replace two with one if they constitute a +common subexpression. Also, if your instruction does have a side effect on +a variable that otherwise appears not to change, the old value of the +variable may be reused later if it happens to be found in a register. + +You can prevent an `asm' instruction from being deleted, moved or combined +by writing the keyword `volatile' after the `asm'. For example: + + #define set_priority(x) \ + asm volatile ("set_priority %1": \ + "=m" (*(char *)0): "g" (x)) + +Note that we have supplied an output operand which is not actually used in +the instruction. This is because `asm' requires at least one output +operand. This requirement exists for internal implementation reasons and +we might be able to relax it in the future. + +In this case output operand has the additional benefit effect of giving the +appearance of writing in memory. As a result, GNU CC will assume that data +previously fetched from memory must be fetched again if needed again later. + This may be desirable if you have not employed the `volatile' keyword on +all the variable declarations that ought to have it.  -File: internals Node: Modifiers, Prev: Class Preferences, Up: Constraints, Next: No Constraints +File: internals, Node: Asm Labels, Prev: Extended Asm, Up: Extensions -Constraint Modifier Characters ------------------------------- +Controlling Names Used in Assembler Code +======================================== -`=' - Means that this operand is written by the instruction, but its previous - value is not used. - -`+' - Means that this operand is both read and written by the instruction. - - When the compiler fixes up the operands to satisfy the constraints, - it needs to know which operands are inputs to the instruction and - which are outputs from it. `=' identifies an output; `+' - identifies an operand that is both input and output; all other operands - are assumed to be input only. - -`%' - Declares the instruction to be commutative for operands 1 and 2. - This means that the compiler may interchange operands 1 and 2 - if that will make the operands fit their constraints. - -`#' - Says that all following characters, up to the next comma, are to be ignored - as a constraint. They are significant only for choosing register preferences. +You can specify the name to be used in the assembler code for a C function +or variable by writing the `asm' keyword after the declarator as follows: + + int foo asm ("myfoo") = 2; + +This specifies that the name to be used for the variable `foo' in the +assembler code should be `myfoo' rather than the usual `_foo'. + +On systems where an underscore is normally prepended to the name of a C +function or variable, this feature allows you to define names for the +linker that do not start with an underscore. + +You cannot use `asm' in this way in a function *definition*; but you can +get the same effect by writing a declaration for the function before its +definition and putting `asm' there, like this: + + extern func () asm ("FUNC"); -`*' - Says that the following character should be ignored when choosing - register preferences. `*' has no effect on the meaning of - the constraint as a constraint. + func (x, y) + int x, y; + ... + + It is up to you to make sure that the assembler names you choose do not +conflict with any other assembler symbols. Also, you must not use a +register name; that would produce completely invalid assembler code. GNU +CC does not as yet have the ability to store static variables in registers. + Perhaps that will be added.  -File: internals Node: No Constraints, Prev: Modifiers, Up: Constraints +File: internals, Node: Bugs, Next: Portability, Prev: Extensions, Up: Top + +Reporting Bugs +************** + +Your bug reports play an essential role in making GNU CC reliable. -Not Using Constraints ---------------------- +Reporting a bug may help you by bringing a solution to your problem, or it +may not. But in any case the important function of a bug report is to help +the entire community by making the next version of GNU CC work better. Bug +reports are your contribution to the maintenance of GNU CC. -Some machines are so clean that operand constraints are not required. For -example, on the Vax, an operand valid in one context is valid in any other -context. On such a machine, every operand constraint would be `"g"', -excepting only operands of "load address" instructions which are -written as if they referred to a memory location's contents but actual -refer to its address. They would have constraint `"p"'. +In order for a bug report to serve its purpose, you must include the +information that makes for fixing the bug. -For such machines, instead of writing `"g"' and `"p"' for all -the constraints, you can choose to write a description with empty constraints. -Then you write `""' for the constraint in every `match_operand'. -Address operands are identified by writing an `address' expression -around the `match_operand', not by their constraints. +* Menu: + +* Criteria: Bug Criteria. Have you really found a bug? +* Reporting: Bug Reporting. How to report a bug effectively. -When the machine description has just empty constraints, certain parts -of compilation are skipped, making the compiler faster.  -File: internals Node: Standard Names, Prev: Constraints, Up: Machine Desc, Next: Dependent Patterns +File: internals, Node: Bug Criteria, Next: Bug Reporting, Prev: Bugs, Up: Bugs -Standard Insn Names -=================== +Have You Found a Bug? +===================== -Here is a table of the instruction names that are meaningful in the RTL -generation pass of the compiler. Giving one of these names to an -instruction pattern tells the RTL generation pass that it can use the -pattern in to accomplish a certain task. - -`movM' - Here M is a two-letter machine mode name, in lower case. This - instruction pattern moves data with that machine mode from operand 1 to - operand 0. For example, `movsi' moves full-word data. - - If operand 0 is a `subreg' with mode M of a register whose - natural mode is wider than M, the effect of this instruction is - to store the specified value in the part of the register that corresponds - to mode M. The effect on the rest of the register is undefined. - -`movstrictM' - Like `movM' except that if operand 0 is a `subreg' - with mode M of a register whose natural mode is wider, - the `movstrictM' instruction is guaranteed not to alter - any of the register except the part which belongs to mode M. - -`addM3' - Add operand 2 and operand 1, storing the result in operand 0. All operands - must have mode M. This can be used even on two-address machines, by - means of constraints requiring operands 1 and 0 to be the same location. - -`subM3' -`mulM3' -`umulM3' -`divM3' -`udivM3' -`modM3' -`umodM3' -`andM3' -`iorM3' -`xorM3' - Similar, for other arithmetic operations. - -`andcbM3' - Bitwise logical-and operand 1 with the complement of operand 2 - and store the result in operand 0. - -`mulhisi3' - Multiply operands 1 and 2, which have mode `HImode', and store - a `SImode' product in operand 0. - -`mulqihi3' -`mulsidi3' - Similar widening-multiplication instructions of other widths. - -`umulqihi3' -`umulhisi3' -`umulsidi3' - Similar widening-multiplication instructions that do unsigned - multiplication. - -`divmodM4' - Signed division that produces both a quotient and a remainder. - Operand 1 is divided by operand 2 to produce a quotient stored - in operand 0 and a remainder stored in operand 3. - -`udivmodM4' - Similar, but does unsigned division. - -`divmodMN4' - Like `divmodM4' except that only the dividend has mode - M; the divisor, quotient and remainder have mode N. - For example, the Vax has a `divmoddisi4' instruction - (but it is omitted from the machine description, because it - is so slow that it is faster to compute remainders by the - circumlocution that the compiler will use if this instruction is - not available). - -`ashlM3' - Arithmetic-shift operand 1 left by a number of bits specified by - operand 2, and store the result in operand 0. Operand 2 has - mode `SImode', not mode M. - -`ashrM3' -`lshlM3' -`lshrM3' -`rotlM3' -`rotrM3' - Other shift and rotate instructions. - -`negM2' - Negate operand 1 and store the result in operand 0. - -`absM2' - Store the absolute value of operand 1 into operand 0. - -`sqrtM2' - Store the square root of operand 1 into operand 0. - -`one_cmplM2' - Store the bitwise-complement of operand 1 into operand 0. - -`cmpM' - Compare operand 0 and operand 1, and set the condition codes. - -`tstM' - Compare operand 0 against zero, and set the condition codes. - -`movstrM' - Block move instruction. The addresses of the destination and source - strings are the first two operands, and both are in mode `Pmode'. - The number of bytes to move is the third operand, in mode M. - -`cmpstrM' - Block compare instruction, with operands like `movstrM' - except that the two memory blocks are compared byte by byte - in lexicographic order. The effect of the instruction is to set - the condition codes. - -`floatMN2' - Convert operand 1 (valid for floating point mode M) to fixed - point mode N and store in operand 0 (which has mode N). - -`fixMN2' - Convert operand 1 (valid for fixed point mode M) to floating - point mode N and store in operand 0 (which has mode N). - -`truncMN' - Truncate operand 1 (valid for mode M) to mode N and - store in operand 0 (which has mode N). Both modes must be fixed - point or both floating point. - -`extendMN' - Sign-extend operand 1 (valid for mode M) to mode N and - store in operand 0 (which has mode N). Both modes must be fixed - point or both floating point. - -`zero_extendMN' - Zero-extend operand 1 (valid for mode M) to mode N and - store in operand 0 (which has mode N). Both modes must be fixed - point. - -`extv' - Extract a bit-field from operand 1 (a register or memory operand), - where operand 2 specifies the width in bits and operand 3 the starting - bit, and store it in operand 0. Operand 0 must have `Simode'. - Operand 1 may have mode `QImode' or `SImode'; often - `SImode' is allowed only for registers. Operands 2 and 3 must be - valid for `SImode'. - - The RTL generation pass generates this instruction only with constants - for operands 2 and 3. - - The bit-field value is sign-extended to a full word integer - before it is stored in operand 0. - -`extzv' - Like `extv' except that the bit-field value is zero-extended. - -`insv' - Store operand 3 (which must be valid for `SImode') into a - bit-field in operand 0, where operand 1 specifies the width in bits - and operand 2 the starting bit. Operand 0 may have mode `QImode' - or `SImode'; often `SImode' is allowed only for registers. - Operands 1 and 2 must be valid for `SImode'. - - The RTL generation pass generates this instruction only with constants - for operands 1 and 2. - -`sCONDM' - Store zero or -1 in the operand (with mode M) according to the - condition codes. Value stored is -1 iff the condition COND is - true. COND is the name of a comparison operation rtx code, such - as `eq', `lt' or `leu'. - -`bCOND' - Conditional branch instruction. Operand 0 is a `label_ref' - that refers to the label to jump to. Jump if the condition codes - meet condition COND. - -`call' - Subroutine call instruction. Operand 1 is the number of arguments - and operand 0 is the function to call. Operand 1 should be a `mem' - rtx whose address is the address of the function. - -`return' - Subroutine return instruction. This instruction pattern name should be - defined only if a single instruction can do all the work of returning - from a function. - -`tablejump' -`caseM' +If you are not sure whether you have found a bug, here are some guidelines: + + * If the compiler gets a fatal signal, for any input whatever, that is a + compiler bug. Reliable compilers never crash. + + * If the compiler produces invalid assembly code, for any input whatever + (except an `asm' statement), that is a compiler bug, unless the + compiler reports errors (not just warnings) which would ordinarily + prevent the assembler from being run. + + * If the compiler produces valid assembly code that does not correctly + execute the input source code, that is a compiler bug. + + However, you must double-check to make sure, because you may have run + into an incompatibility between GNU C and traditional C (*Note + Incompatibilities::.). These incompatibilities might be considered + bugs, but they are inescapable consequences of valuable features. + + Or you may have a program whose behavior is undefined, which happened + by chance to give the desired results with another C compiler. + + For example, in many nonoptimizing compilers, you can write `x;' at + the end of a function instead of `return x;', with the same results. + But the value of the function is undefined if `return' is omitted; it + is not a bug when GNU CC produces different results. + + Problems often result from expressions with two increment operators, + as in `f (*p++, *p++)'. Your previous compiler might have interpreted + that expression the way you intended; GNU CC might interpret it + another way; neither compiler is wrong. + + After you have localized the error to a single source line, it should + be easy to check for these things. If your program is correct and + well defined, you have found a compiler bug. + + * If the compiler produces an error message for valid input, that is a + compiler bug. + + Note that the following is not valid input, and the error message for + it is not a bug: + + int foo (char); + + int + foo (x) + char x; + { ... } + + The prototype says to pass a `char', while the definition says to pass + an `int' and treat the value as a `char'. This is what the ANSI + standard says, and it makes sense. + + * If the compiler does not produce an error message for invalid input, + that is a compiler bug. However, you should note that your idea of + ``invalid input'' might be my idea of ``an extension'' or ``support + for traditional practice''. + + * If you are an experienced user of C compilers, your suggestions for + improvement of GNU CC are welcome in any case. + + +File: internals, Node: Bug Reporting, Prev: Bug Criteria, Up: Bugs + +How to Report Bugs +================== + +Send bug reports for GNU C to one of these addresses: + + bug-gcc@prep.ai.mit.edu + {ucbvax|mit-eddie|uunet}!prep.ai.mit.edu!bug-gcc + +As a last resort, snail them to: + + GNU Compiler Bugs + 545 Tech Sq + Cambridge, MA 02139 + +The fundamental principle of reporting bugs usefully is this: *report all +the facts*. If you are not sure whether to mention a fact or leave it out, +mention it! + +Often people omit facts because they think they know what causes the +problem and they conclude that some details don't matter. Thus, you might +assume that the name of the variable you use in an example does not matter. + Well, probably it doesn't, but one cannot be sure. Perhaps the bug is a +stray memory reference which happens to fetch from the location where that +name is stored in memory; perhaps, if the name were different, the contents +of that location would fool the compiler into doing the right thing despite +the bug. Play it safe and give an exact example. + +If you want to enable me to fix the bug, you should include all these things: + + * The version of GNU CC. You can get this by running it with the `-v' + option. + + Without this, I won't know whether there is any point in looking for + the bug in the current version of GNU CC. + + * A complete input file that will reproduce the bug. If the bug is in + the C preprocessor, send me a source file and any header files that it + requires. If the bug is in the compiler proper (`cc1'), run your + source file through the C preprocessor by doing `gcc -E SOURCEFILE > + OUTFILE', then include the contents of OUTFILE in the bug report. + (Any `-I', `-D' or `-U' options that you used in actual compilation + should also be used when doing this.) + + A single statement is not enough of an example. In order to compile + it, it must be embedded in a function definition; and the bug might + depend on the details of how this is done. + + Without a real example I can compile, all I can do about your bug + report is wish you luck. It would be futile to try to guess how to + provoke the bug. For example, bugs in register allocation and + reloading frequently depend on every little detail of the function + they happen in. + + * The command arguments you gave GNU CC to compile that example and + observe the bug. For example, did you use `-O'? To guarantee you + won't omit something important, list them all. + + If I were to try to guess the arguments, I would probably guess wrong + and then I would not encounter the bug. + + * The names of the files that you used for `tm.h' and `md' when you + installed the compiler. + + * The type of machine you are using, and the operating system name and + version number. + + * A description of what behavior you observe that you believe is + incorrect. For example, ``It gets a fatal signal,'' or, ``There is an + incorrect assembler instruction in the output.'' + + Of course, if the bug is that the compiler gets a fatal signal, then I + will certainly notice it. But if the bug is incorrect output, I might + not notice unless it is glaringly wrong. I won't study all the + assembler code from a 50-line C program just on the off chance that it + might be wrong. + + Even if the problem you experience is a fatal signal, you should still + say so explicitly. Suppose something strange is going on, such as, + your copy of the compiler is out of synch, or you have encountered a + bug in the C library on your system. (This has happened!) Your copy + might crash and mine would not. If you told me to expect a crash, + then when mine fails to crash, I would know that the bug was not + happening for me. If you had not told me to expect a crash, then I + would not be able to draw any conclusion from my observations. + + In cases where GNU CC generates incorrect code, if you send me a small + complete sample program I will find the error myself by running the + program under a debugger. If you send me a large example or a part of + a larger program, I cannot do this; you must debug the compiled + program and narrow the problem down to one source line. Tell me which + source line it is, and what you believe is incorrect about the code + generated for that line. + + * If you send me examples of output from GNU CC, please use `-g' when + you make them. The debugging information includes source line numbers + which are essential for correlating the output with the input. + +Here are some things that are not necessary: + + * A description of the envelope of the bug. + + Often people who encounter a bug spend a lot of time investigating + which changes to the input file will make the bug go away and which + changes will not affect it. + + This is often time consuming and not very useful, because the way I + will find the bug is by running a single example under the debugger + with breakpoints, not by pure deduction from a series of examples. + + Of course, it can't hurt if you can find a simpler example that + triggers the same bug. Errors in the output will be easier to spot, + running under the debugger will take less time, etc. An easy way to + simplify an example is to delete all the function definitions except + the one where the bug occurs. Those earlier in the file may be + replaced by external declarations. + + However, simplification is not necessary; if you don't want to do + this, report the bug anyway. + + * A patch for the bug. + + A patch for the bug does help me if it is a good one. But don't omit + the necessary information, such as the test case, because I might see + problems with your patch and decide to fix the problem another way. + + Sometimes with a program as complicated as GNU CC it is very hard to + construct an example that will make the program go through a certain + point in the code. If you don't send me the example, I won't be able + to verify that the bug is fixed. + + * A guess about what the bug is or what it depends on. + + Such guesses are usually wrong. Even I can't guess right about such + things without using the debugger to find the facts. They also don't + serve a useful purpose. + + +File: internals, Node: Portability, Next: Interface, Prev: Bugs, Up: Top + +GNU CC and Portability +********************** + +The main goal of GNU CC was to make a good, fast compiler for machines in +the class that the GNU system aims to run on: 32-bit machines that address +8-bit bytes and have several general registers. Elegance, theoretical +power and simplicity are only secondary. + +GNU CC gets most of the information about the target machine from a machine +description which gives an algebraic formula for each of the machine's +instructions. This is a very clean way to describe the target. But when +the compiler needs information that is difficult to express in this +fashion, I have not hesitated to define an ad-hoc parameter to the machine +description. The purpose of portability is to reduce the total work needed +on the compiler; it was not of interest for its own sake. + +GNU CC does not contain machine dependent code, but it does contain code +that depends on machine parameters such as endianness (whether the most +significant byte has the highest or lowest address of the bytes in a word) +and the availability of autoincrement addressing. In the RTL-generation +pass, it is often necessary to have multiple strategies for generating code +for a particular kind of syntax tree, strategies that are usable for +different combinations of parameters. Often I have not tried to address +all possible cases, but only the common ones or only the ones that I have +encountered. As a result, a new target may require additional strategies. +You will know if this happens because the compiler will call `abort'. +Fortunately, the new strategies can be added in a machine-independent +fashion, and will affect only the target machines that need them. + + +File: internals, Node: Interface, Next: Passes, Prev: Portability, Up: Top + +Interfacing to GNU CC Output +**************************** + +GNU CC is normally configured to use the same function calling convention +normally in use on the target system. This is done with the +machine-description macros described (*Note Machine Macros::.). + +However, returning of structure and union values is done differently. As a +result, functions compiled with PCC returning such types cannot be called +from code compiled with GNU CC, and vice versa. This usually does not +cause trouble because the Unix library routines don't return structures and +unions. + +Structures and unions that are 1, 2, 4 or 8 bytes long are returned in the +same registers used for `int' or `double' return values. (GNU CC typically +allocates variables of such types in registers also.) Structures and +unions of other sizes are returned by storing them into an address passed +by the caller in a register. This method is faster than the one normally +used by PCC and is also reentrant. The register used for passing the +address is specified by the machine-description macro `STRUCT_VALUE_REGNUM'. + +GNU CC always passes arguments on the stack. At some point it will be +extended to pass arguments in registers, for machines which use that as the +standard calling convention. This will make it possible to use such a +convention on other machines as well. However, that would render it +completely incompatible with PCC. We will probably do this once we have a +complete GNU system so we can compile the libraries with GNU CC. + +If you use `longjmp', beware of automatic variables. ANSI C says that +automatic variables that are not declared `volatile' have undefined values +after a `longjmp'. And this is all GNU CC promises to do, because it is +very difficult to restore register variables correctly, and one of GNU CC's +features is that it can put variables in registers without your asking it to. + +If you want a variable to be unaltered by `longjmp', and you don't want to +write `volatile' because old C compilers don't accept it, just take the +address of the variable. If a variable's address is ever taken, even if +just to compute it and ignore it, then the variable cannot go in a register: + + { + int careful; + &careful; + ... + } + +Code compiled with GNU CC may call certain library routines. The routines +needed on the Vax and 68000 are in the file `gnulib.c'. You must compile +this file with the standard C compiler, not with GNU CC, and then link it +with each program you compile with GNU CC. (In actuality, many programs +will not need it.) The usual function call interface is used for calling +the library routines. Some standard parts of the C library, such as +`bcopy', are also called automatically.  \ No newline at end of file