--- gcc/gcc.info-14 2018/04/24 18:17:47 1.1.1.7 +++ gcc/gcc.info-14 2018/04/24 18:24:08 1.1.1.8 @@ -3,11 +3,11 @@ file gcc.texi. This file documents the use and the internals of the GNU compiler. - Published by the Free Software Foundation 675 Massachusetts Avenue -Cambridge, MA 02139 USA + Published by the Free Software Foundation 59 Temple Place - Suite 330 +Boston, MA 02111-1307 USA - Copyright (C) 1988, 1989, 1992, 1993, 1994 Free Software Foundation, -Inc. + Copyright (C) 1988, 1989, 1992, 1993, 1994, 1995 Free Software +Foundation, Inc. Permission is granted to make and distribute verbatim copies of this manual provided the copyright notice and this permission notice are @@ -30,1071 +30,941 @@ translations approved by the Free Softwa original English.  -File: gcc.info, Node: Insns, Next: Calls, Prev: Assembler, Up: RTL +File: gcc.info, Node: Machine Modes, Next: Constants, Prev: Flags, Up: RTL -Insns -===== +Machine Modes +============= - The RTL representation of the code for a function is a doubly-linked -chain of objects called "insns". Insns are expressions with special -codes that are used for no other purpose. Some insns are actual -instructions; others represent dispatch tables for `switch' statements; -others represent labels to jump to or various sorts of declarative -information. - - In addition to its own specific data, each insn must have a unique -id-number that distinguishes it from all other insns in the current -function (after delayed branch scheduling, copies of an insn with the -same id-number may be present in multiple places in a function, but -these copies will always be identical and will only appear inside a -`sequence'), and chain pointers to the preceding and following insns. -These three fields occupy the same position in every insn, independent -of the expression code of the insn. They could be accessed with `XEXP' -and `XINT', but instead three special macros are always used: - -`INSN_UID (I)' - Accesses the unique id of insn I. - -`PREV_INSN (I)' - Accesses the chain pointer to the insn preceding I. If I is the - first insn, this is a null pointer. - -`NEXT_INSN (I)' - Accesses the chain pointer to the insn following I. If I is the - last insn, this is a null pointer. - - The first insn in the chain is obtained by calling `get_insns'; the -last insn is the result of calling `get_last_insn'. Within the chain -delimited by these insns, the `NEXT_INSN' and `PREV_INSN' pointers must -always correspond: if INSN is not the first insn, - - NEXT_INSN (PREV_INSN (INSN)) == INSN - -is always true and if INSN is not the last insn, - - PREV_INSN (NEXT_INSN (INSN)) == INSN - -is always true. - - After delay slot scheduling, some of the insns in the chain might be -`sequence' expressions, which contain a vector of insns. The value of -`NEXT_INSN' in all but the last of these insns is the next insn in the -vector; the value of `NEXT_INSN' of the last insn in the vector is the -same as the value of `NEXT_INSN' for the `sequence' in which it is -contained. Similar rules apply for `PREV_INSN'. - - This means that the above invariants are not necessarily true for -insns inside `sequence' expressions. Specifically, if INSN is the -first insn in a `sequence', `NEXT_INSN (PREV_INSN (INSN))' is the insn -containing the `sequence' expression, as is the value of `PREV_INSN -(NEXT_INSN (INSN))' is INSN is the last insn in the `sequence' -expression. You can use these expressions to find the containing -`sequence' expression. - - Every insn has one of the following six expression codes: - -`insn' - The expression code `insn' is used for instructions that do not - jump and do not do function calls. `sequence' expressions are - always contained in insns with code `insn' even if one of those - insns should jump or do function calls. - - Insns with code `insn' have four additional fields beyond the three - mandatory ones listed above. These four are described in a table - below. - -`jump_insn' - The expression code `jump_insn' is used for instructions that may - jump (or, more generally, may contain `label_ref' expressions). If - there is an instruction to return from the current function, it is - recorded as a `jump_insn'. - - `jump_insn' insns have the same extra fields as `insn' insns, - accessed in the same way and in addition contain a field - `JUMP_LABEL' which is defined once jump optimization has completed. - - For simple conditional and unconditional jumps, this field - contains the `code_label' to which this insn will (possibly - conditionally) branch. In a more complex jump, `JUMP_LABEL' - records one of the labels that the insn refers to; the only way to - find the others is to scan the entire body of the insn. - - Return insns count as jumps, but since they do not refer to any - labels, they have zero in the `JUMP_LABEL' field. - -`call_insn' - The expression code `call_insn' is used for instructions that may - do function calls. It is important to distinguish these - instructions because they imply that certain registers and memory - locations may be altered unpredictably. - - `call_insn' insns have the same extra fields as `insn' insns, - accessed in the same way and in addition contain a field - `CALL_INSN_FUNCTION_USAGE', which contains a list (chain of - `expr_list' expressions) containing `use' and `clobber' - expressions that denote hard registers used or clobbered by the - called function. A register specified in a `clobber' in this list - is modified *after* the execution of the `call_insn', while a - register in a `clobber' in the body of the `call_insn' is - clobbered before the insn completes execution. `clobber' - expressions in this list augment registers specified in - `CALL_USED_REGISTERS' (*note Register Basics::.). - -`code_label' - A `code_label' insn represents a label that a jump insn can jump - to. It contains two special fields of data in addition to the - three standard ones. `CODE_LABEL_NUMBER' is used to hold the - "label number", a number that identifies this label uniquely among - all the labels in the compilation (not just in the current - function). Ultimately, the label is represented in the assembler - output as an assembler label, usually of the form `LN' where N is - the label number. - - When a `code_label' appears in an RTL expression, it normally - appears within a `label_ref' which represents the address of the - label, as a number. - - The field `LABEL_NUSES' is only defined once the jump optimization - phase is completed and contains the number of times this label is - referenced in the current function. - -`barrier' - Barriers are placed in the instruction stream when control cannot - flow past them. They are placed after unconditional jump - instructions to indicate that the jumps are unconditional and - after calls to `volatile' functions, which do not return (e.g., - `exit'). They contain no information beyond the three standard - fields. - -`note' - `note' insns are used to represent additional debugging and - declarative information. They contain two nonstandard fields, an - integer which is accessed with the macro `NOTE_LINE_NUMBER' and a - string accessed with `NOTE_SOURCE_FILE'. - - If `NOTE_LINE_NUMBER' is positive, the note represents the - position of a source line and `NOTE_SOURCE_FILE' is the source - file name that the line came from. These notes control generation - of line number data in the assembler output. - - Otherwise, `NOTE_LINE_NUMBER' is not really a line number but a - code with one of the following values (and `NOTE_SOURCE_FILE' must - contain a null pointer): - - `NOTE_INSN_DELETED' - Such a note is completely ignorable. Some passes of the - compiler delete insns by altering them into notes of this - kind. - - `NOTE_INSN_BLOCK_BEG' - `NOTE_INSN_BLOCK_END' - These types of notes indicate the position of the beginning - and end of a level of scoping of variable names. They - control the output of debugging information. - - `NOTE_INSN_LOOP_BEG' - `NOTE_INSN_LOOP_END' - These types of notes indicate the position of the beginning - and end of a `while' or `for' loop. They enable the loop - optimizer to find loops quickly. - - `NOTE_INSN_LOOP_CONT' - Appears at the place in a loop that `continue' statements - jump to. - - `NOTE_INSN_LOOP_VTOP' - This note indicates the place in a loop where the exit test - begins for those loops in which the exit test has been - duplicated. This position becomes another virtual start of - the loop when considering loop invariants. - - `NOTE_INSN_FUNCTION_END' - Appears near the end of the function body, just before the - label that `return' statements jump to (on machine where a - single instruction does not suffice for returning). This - note may be deleted by jump optimization. - - `NOTE_INSN_SETJMP' - Appears following each call to `setjmp' or a related function. - - These codes are printed symbolically when they appear in debugging - dumps. - - The machine mode of an insn is normally `VOIDmode', but some phases -use the mode for various purposes; for example, the reload pass sets it -to `HImode' if the insn needs reloading but not register elimination -and `QImode' if both are required. The common subexpression -elimination pass sets the mode of an insn to `QImode' when it is the -first insn in a block that has already been processed. - - Here is a table of the extra fields of `insn', `jump_insn' and -`call_insn' insns: - -`PATTERN (I)' - An expression for the side effect performed by this insn. This - must be one of the following codes: `set', `call', `use', - `clobber', `return', `asm_input', `asm_output', `addr_vec', - `addr_diff_vec', `trap_if', `unspec', `unspec_volatile', - `parallel', or `sequence'. If it is a `parallel', each element of - the `parallel' must be one these codes, except that `parallel' - expressions cannot be nested and `addr_vec' and `addr_diff_vec' - are not permitted inside a `parallel' expression. - -`INSN_CODE (I)' - An integer that says which pattern in the machine description - matches this insn, or -1 if the matching has not yet been - attempted. - - Such matching is never attempted and this field remains -1 on an - insn whose pattern consists of a single `use', `clobber', - `asm_input', `addr_vec' or `addr_diff_vec' expression. - - Matching is also never attempted on insns that result from an `asm' - statement. These contain at least one `asm_operands' expression. - The function `asm_noperands' returns a non-negative value for such - insns. - - In the debugging output, this field is printed as a number - followed by a symbolic representation that locates the pattern in - the `md' file as some small positive or negative offset from a - named pattern. - -`LOG_LINKS (I)' - A list (chain of `insn_list' expressions) giving information about - dependencies between instructions within a basic block. Neither a - jump nor a label may come between the related insns. - -`REG_NOTES (I)' - A list (chain of `expr_list' and `insn_list' expressions) giving - miscellaneous information about the insn. It is often information - pertaining to the registers used in this insn. - - The `LOG_LINKS' field of an insn is a chain of `insn_list' -expressions. Each of these has two operands: the first is an insn, and -the second is another `insn_list' expression (the next one in the -chain). The last `insn_list' in the chain has a null pointer as second -operand. The significant thing about the chain is which insns appear -in it (as first operands of `insn_list' expressions). Their order is -not significant. - - This list is originally set up by the flow analysis pass; it is a -null pointer until then. Flow only adds links for those data -dependencies which can be used for instruction combination. For each -insn, the flow analysis pass adds a link to insns which store into -registers values that are used for the first time in this insn. The -instruction scheduling pass adds extra links so that every dependence -will be represented. Links represent data dependencies, -antidependencies and output dependencies; the machine mode of the link -distinguishes these three types: antidependencies have mode -`REG_DEP_ANTI', output dependencies have mode `REG_DEP_OUTPUT', and -data dependencies have mode `VOIDmode'. - - The `REG_NOTES' field of an insn is a chain similar to the -`LOG_LINKS' field but it includes `expr_list' expressions in addition -to `insn_list' expressions. There are several kinds of register notes, -which are distinguished by the machine mode, which in a register note -is really understood as being an `enum reg_note'. The first operand OP -of the note is data whose meaning depends on the kind of note. - - The macro `REG_NOTE_KIND (X)' returns the kind of register note. -Its counterpart, the macro `PUT_REG_NOTE_KIND (X, NEWKIND)' sets the -register note type of X to be NEWKIND. - - Register notes are of three classes: They may say something about an -input to an insn, they may say something about an output of an insn, or -they may create a linkage between two insns. There are also a set of -values that are only used in `LOG_LINKS'. - - These register notes annotate inputs to an insn: - -`REG_DEAD' - The value in OP dies in this insn; that is to say, altering the - value immediately after this insn would not affect the future - behavior of the program. - - This does not necessarily mean that the register OP has no useful - value after this insn since it may also be an output of the insn. - In such a case, however, a `REG_DEAD' note would be redundant and - is usually not present until after the reload pass, but no code - relies on this fact. - -`REG_INC' - The register OP is incremented (or decremented; at this level - there is no distinction) by an embedded side effect inside this - insn. This means it appears in a `post_inc', `pre_inc', - `post_dec' or `pre_dec' expression. - -`REG_NONNEG' - The register OP is known to have a nonnegative value when this - insn is reached. This is used so that decrement and branch until - zero instructions, such as the m68k dbra, can be matched. - - The `REG_NONNEG' note is added to insns only if the machine - description has a `decrement_and_branch_until_zero' pattern. - -`REG_NO_CONFLICT' - This insn does not cause a conflict between OP and the item being - set by this insn even though it might appear that it does. In - other words, if the destination register and OP could otherwise be - assigned the same register, this insn does not prevent that - assignment. - - Insns with this note are usually part of a block that begins with a - `clobber' insn specifying a multi-word pseudo register (which will - be the output of the block), a group of insns that each set one - word of the value and have the `REG_NO_CONFLICT' note attached, - and a final insn that copies the output to itself with an attached - `REG_EQUAL' note giving the expression being computed. This block - is encapsulated with `REG_LIBCALL' and `REG_RETVAL' notes on the - first and last insns, respectively. - -`REG_LABEL' - This insn uses OP, a `code_label', but is not a `jump_insn'. The - presence of this note allows jump optimization to be aware that OP - is, in fact, being used. - - The following notes describe attributes of outputs of an insn: - -`REG_EQUIV' -`REG_EQUAL' - This note is only valid on an insn that sets only one register and - indicates that that register will be equal to OP at run time; the - scope of this equivalence differs between the two types of notes. - The value which the insn explicitly copies into the register may - look different from OP, but they will be equal at run time. If the - output of the single `set' is a `strict_low_part' expression, the - note refers to the register that is contained in `SUBREG_REG' of - the `subreg' expression. - - For `REG_EQUIV', the register is equivalent to OP throughout the - entire function, and could validly be replaced in all its - occurrences by OP. ("Validly" here refers to the data flow of the - program; simple replacement may make some insns invalid.) For - example, when a constant is loaded into a register that is never - assigned any other value, this kind of note is used. - - When a parameter is copied into a pseudo-register at entry to a - function, a note of this kind records that the register is - equivalent to the stack slot where the parameter was passed. - Although in this case the register may be set by other insns, it - is still valid to replace the register by the stack slot - throughout the function. - - In the case of `REG_EQUAL', the register that is set by this insn - will be equal to OP at run time at the end of this insn but not - necessarily elsewhere in the function. In this case, OP is - typically an arithmetic expression. For example, when a sequence - of insns such as a library call is used to perform an arithmetic - operation, this kind of note is attached to the insn that produces - or copies the final value. - - These two notes are used in different ways by the compiler passes. - `REG_EQUAL' is used by passes prior to register allocation (such as - common subexpression elimination and loop optimization) to tell - them how to think of that value. `REG_EQUIV' notes are used by - register allocation to indicate that there is an available - substitute expression (either a constant or a `mem' expression for - the location of a parameter on the stack) that may be used in - place of a register if insufficient registers are available. - - Except for stack homes for parameters, which are indicated by a - `REG_EQUIV' note and are not useful to the early optimization - passes and pseudo registers that are equivalent to a memory - location throughout there entire life, which is not detected until - later in the compilation, all equivalences are initially indicated - by an attached `REG_EQUAL' note. In the early stages of register - allocation, a `REG_EQUAL' note is changed into a `REG_EQUIV' note - if OP is a constant and the insn represents the only set of its - destination register. - - Thus, compiler passes prior to register allocation need only check - for `REG_EQUAL' notes and passes subsequent to register allocation - need only check for `REG_EQUIV' notes. - -`REG_UNUSED' - The register OP being set by this insn will not be used in a - subsequent insn. This differs from a `REG_DEAD' note, which - indicates that the value in an input will not be used subsequently. - These two notes are independent; both may be present for the same - register. - -`REG_WAS_0' - The single output of this insn contained zero before this insn. - OP is the insn that set it to zero. You can rely on this note if - it is present and OP has not been deleted or turned into a `note'; - its absence implies nothing. - - These notes describe linkages between insns. They occur in pairs: -one insn has one of a pair of notes that points to a second insn, which -has the inverse note pointing back to the first insn. - -`REG_RETVAL' - This insn copies the value of a multi-insn sequence (for example, a - library call), and OP is the first insn of the sequence (for a - library call, the first insn that was generated to set up the - arguments for the library call). - - Loop optimization uses this note to treat such a sequence as a - single operation for code motion purposes and flow analysis uses - this note to delete such sequences whose results are dead. - - A `REG_EQUAL' note will also usually be attached to this insn to - provide the expression being computed by the sequence. - -`REG_LIBCALL' - This is the inverse of `REG_RETVAL': it is placed on the first - insn of a multi-insn sequence, and it points to the last one. - -`REG_CC_SETTER' -`REG_CC_USER' - On machines that use `cc0', the insns which set and use `cc0' set - and use `cc0' are adjacent. However, when branch delay slot - filling is done, this may no longer be true. In this case a - `REG_CC_USER' note will be placed on the insn setting `cc0' to - point to the insn using `cc0' and a `REG_CC_SETTER' note will be - placed on the insn using `cc0' to point to the insn setting `cc0'. - - These values are only used in the `LOG_LINKS' field, and indicate -the type of dependency that each link represents. Links which indicate -a data dependence (a read after write dependence) do not use any code, -they simply have mode `VOIDmode', and are printed without any -descriptive text. - -`REG_DEP_ANTI' - This indicates an anti dependence (a write after read dependence). - -`REG_DEP_OUTPUT' - This indicates an output dependence (a write after write - dependence). - - For convenience, the machine mode in an `insn_list' or `expr_list' -is printed using these symbolic codes in debugging dumps. - - The only difference between the expression codes `insn_list' and -`expr_list' is that the first operand of an `insn_list' is assumed to -be an insn and is printed in debugging dumps as the insn's unique id; -the first operand of an `expr_list' is printed in the ordinary way as -an expression. + A machine mode describes a size of data object and the +representation used for it. In the C code, machine modes are +represented by an enumeration type, `enum machine_mode', defined in +`machmode.def'. Each RTL expression has room for a machine mode and so +do certain kinds of tree expressions (declarations and types, to be +precise). + + In debugging dumps and machine descriptions, the machine mode of an +RTL expression is written after the expression code with a colon to +separate them. The letters `mode' which appear at the end of each +machine mode name are omitted. For example, `(reg:SI 38)' is a `reg' +expression with machine mode `SImode'. If the mode is `VOIDmode', it +is not written at all. + + Here is a table of machine modes. The term "byte" below refers to an +object of `BITS_PER_UNIT' bits (*note Storage Layout::.). + +`QImode' + "Quarter-Integer" mode represents a single byte treated as an + integer. + +`HImode' + "Half-Integer" mode represents a two-byte integer. + +`PSImode' + "Partial Single Integer" mode represents an integer which occupies + four bytes but which doesn't really use all four. On some + machines, this is the right mode to use for pointers. + +`SImode' + "Single Integer" mode represents a four-byte integer. + +`PDImode' + "Partial Double Integer" mode represents an integer which occupies + eight bytes but which doesn't really use all eight. On some + machines, this is the right mode to use for certain pointers. + +`DImode' + "Double Integer" mode represents an eight-byte integer. + +`TImode' + "Tetra Integer" (?) mode represents a sixteen-byte integer. + +`SFmode' + "Single Floating" mode represents a single-precision (four byte) + floating point number. + +`DFmode' + "Double Floating" mode represents a double-precision (eight byte) + floating point number. + +`XFmode' + "Extended Floating" mode represents a triple-precision (twelve + byte) floating point number. This mode is used for IEEE extended + floating point. On some systems not all bits within these bytes + will actually be used. + +`TFmode' + "Tetra Floating" mode represents a quadruple-precision (sixteen + byte) floating point number. + +`CCmode' + "Condition Code" mode represents the value of a condition code, + which is a machine-specific set of bits used to represent the + result of a comparison operation. Other machine-specific modes + may also be used for the condition code. These modes are not used + on machines that use `cc0' (see *note Condition Code::.). + +`BLKmode' + "Block" mode represents values that are aggregates to which none of + the other modes apply. In RTL, only memory references can have + this mode, and only if they appear in string-move or vector + instructions. On machines which have no such instructions, + `BLKmode' will not appear in RTL. + +`VOIDmode' + Void mode means the absence of a mode or an unspecified mode. For + example, RTL expressions of code `const_int' have mode `VOIDmode' + because they can be taken to have whatever mode the context + requires. In debugging dumps of RTL, `VOIDmode' is expressed by + the absence of any mode. + +`SCmode, DCmode, XCmode, TCmode' + These modes stand for a complex number represented as a pair of + floating point values. The floating point values are in `SFmode', + `DFmode', `XFmode', and `TFmode', respectively. + +`CQImode, CHImode, CSImode, CDImode, CTImode, COImode' + These modes stand for a complex number represented as a pair of + integer values. The integer values are in `QImode', `HImode', + `SImode', `DImode', `TImode', and `OImode', respectively. + + The machine description defines `Pmode' as a C macro which expands +into the machine mode used for addresses. Normally this is the mode +whose size is `BITS_PER_WORD', `SImode' on 32-bit machines. + + The only modes which a machine description must support are +`QImode', and the modes corresponding to `BITS_PER_WORD', +`FLOAT_TYPE_SIZE' and `DOUBLE_TYPE_SIZE'. The compiler will attempt to +use `DImode' for 8-byte structures and unions, but this can be +prevented by overriding the definition of `MAX_FIXED_MODE_SIZE'. +Alternatively, you can have the compiler use `TImode' for 16-byte +structures and unions. Likewise, you can arrange for the C type `short +int' to avoid using `HImode'. + + Very few explicit references to machine modes remain in the compiler +and these few references will soon be removed. Instead, the machine +modes are divided into mode classes. These are represented by the +enumeration type `enum mode_class' defined in `machmode.h'. The +possible mode classes are: + +`MODE_INT' + Integer modes. By default these are `QImode', `HImode', `SImode', + `DImode', and `TImode'. + +`MODE_PARTIAL_INT' + The "partial integer" modes, `PSImode' and `PDImode'. + +`MODE_FLOAT' + floating point modes. By default these are `SFmode', `DFmode', + `XFmode' and `TFmode'. + +`MODE_COMPLEX_INT' + Complex integer modes. (These are not currently implemented). + +`MODE_COMPLEX_FLOAT' + Complex floating point modes. By default these are `SCmode', + `DCmode', `XCmode', and `TCmode'. + +`MODE_FUNCTION' + Algol or Pascal function variables including a static chain. + (These are not currently implemented). + +`MODE_CC' + Modes representing condition code values. These are `CCmode' plus + any modes listed in the `EXTRA_CC_MODES' macro. *Note Jump + Patterns::, also see *Note Condition Code::. + +`MODE_RANDOM' + This is a catchall mode class for modes which don't fit into the + above classes. Currently `VOIDmode' and `BLKmode' are in + `MODE_RANDOM'. + + Here are some C macros that relate to machine modes: + +`GET_MODE (X)' + Returns the machine mode of the RTX X. + +`PUT_MODE (X, NEWMODE)' + Alters the machine mode of the RTX X to be NEWMODE. + +`NUM_MACHINE_MODES' + Stands for the number of machine modes available on the target + machine. This is one greater than the largest numeric value of any + machine mode. + +`GET_MODE_NAME (M)' + Returns the name of mode M as a string. + +`GET_MODE_CLASS (M)' + Returns the mode class of mode M. + +`GET_MODE_WIDER_MODE (M)' + Returns the next wider natural mode. For example, the expression + `GET_MODE_WIDER_MODE (QImode)' returns `HImode'. + +`GET_MODE_SIZE (M)' + Returns the size in bytes of a datum of mode M. + +`GET_MODE_BITSIZE (M)' + Returns the size in bits of a datum of mode M. + +`GET_MODE_MASK (M)' + Returns a bitmask containing 1 for all bits in a word that fit + within mode M. This macro can only be used for modes whose + bitsize is less than or equal to `HOST_BITS_PER_INT'. + +`GET_MODE_ALIGNMENT (M))' + Return the required alignment, in bits, for an object of mode M. + +`GET_MODE_UNIT_SIZE (M)' + Returns the size in bytes of the subunits of a datum of mode M. + This is the same as `GET_MODE_SIZE' except in the case of complex + modes. For them, the unit size is the size of the real or + imaginary part. + +`GET_MODE_NUNITS (M)' + Returns the number of units contained in a mode, i.e., + `GET_MODE_SIZE' divided by `GET_MODE_UNIT_SIZE'. + +`GET_CLASS_NARROWEST_MODE (C)' + Returns the narrowest mode in mode class C. + + The global variables `byte_mode' and `word_mode' contain modes whose +classes are `MODE_INT' and whose bitsizes are either `BITS_PER_UNIT' or +`BITS_PER_WORD', respectively. On 32-bit machines, these are `QImode' +and `SImode', respectively.  -File: gcc.info, Node: Calls, Next: Sharing, Prev: Insns, Up: RTL +File: gcc.info, Node: Constants, Next: Regs and Memory, Prev: Machine Modes, Up: RTL -RTL Representation of Function-Call Insns -========================================= +Constant Expression Types +========================= - Insns that call subroutines have the RTL expression code `call_insn'. -These insns must satisfy special rules, and their bodies must use a -special RTL expression code, `call'. - - A `call' expression has two operands, as follows: - - (call (mem:FM ADDR) NBYTES) - -Here NBYTES is an operand that represents the number of bytes of -argument data being passed to the subroutine, FM is a machine mode -(which must equal as the definition of the `FUNCTION_MODE' macro in the -machine description) and ADDR represents the address of the subroutine. - - For a subroutine that returns no value, the `call' expression as -shown above is the entire body of the insn, except that the insn might -also contain `use' or `clobber' expressions. - - For a subroutine that returns a value whose mode is not `BLKmode', -the value is returned in a hard register. If this register's number is -R, then the body of the call insn looks like this: - - (set (reg:M R) - (call (mem:FM ADDR) NBYTES)) - -This RTL expression makes it clear (to the optimizer passes) that the -appropriate register receives a useful value in this insn. - - When a subroutine returns a `BLKmode' value, it is handled by -passing to the subroutine the address of a place to store the value. -So the call insn itself does not "return" any value, and it has the -same RTL form as a call that returns nothing. - - On some machines, the call instruction itself clobbers some register, -for example to contain the return address. `call_insn' insns on these -machines should have a body which is a `parallel' that contains both -the `call' expression and `clobber' expressions that indicate which -registers are destroyed. Similarly, if the call instruction requires -some register other than the stack pointer that is not explicitly -mentioned it its RTL, a `use' subexpression should mention that -register. - - Functions that are called are assumed to modify all registers listed -in the configuration macro `CALL_USED_REGISTERS' (*note Register -Basics::.) and, with the exception of `const' functions and library -calls, to modify all of memory. - - Insns containing just `use' expressions directly precede the -`call_insn' insn to indicate which registers contain inputs to the -function. Similarly, if registers other than those in -`CALL_USED_REGISTERS' are clobbered by the called function, insns -containing a single `clobber' follow immediately after the call to -indicate which registers. + The simplest RTL expressions are those that represent constant +values. - -File: gcc.info, Node: Sharing, Next: Reading RTL, Prev: Calls, Up: RTL - -Structure Sharing Assumptions -============================= +`(const_int I)' + This type of expression represents the integer value I. I is + customarily accessed with the macro `INTVAL' as in `INTVAL (EXP)', + which is equivalent to `XWINT (EXP, 0)'. + + There is only one expression object for the integer value zero; it + is the value of the variable `const0_rtx'. Likewise, the only + expression for integer value one is found in `const1_rtx', the only + expression for integer value two is found in `const2_rtx', and the + only expression for integer value negative one is found in + `constm1_rtx'. Any attempt to create an expression of code + `const_int' and value zero, one, two or negative one will return + `const0_rtx', `const1_rtx', `const2_rtx' or `constm1_rtx' as + appropriate. + + Similarly, there is only one object for the integer whose value is + `STORE_FLAG_VALUE'. It is found in `const_true_rtx'. If + `STORE_FLAG_VALUE' is one, `const_true_rtx' and `const1_rtx' will + point to the same object. If `STORE_FLAG_VALUE' is -1, + `const_true_rtx' and `constm1_rtx' will point to the same object. + +`(const_double:M ADDR I0 I1 ...)' + Represents either a floating-point constant of mode M or an + integer constant too large to fit into `HOST_BITS_PER_WIDE_INT' + bits but small enough to fit within twice that number of bits (GNU + CC does not provide a mechanism to represent even larger + constants). In the latter case, M will be `VOIDmode'. + + ADDR is used to contain the `mem' expression that corresponds to + the location in memory that at which the constant can be found. If + it has not been allocated a memory location, but is on the chain + of all `const_double' expressions in this compilation (maintained + using an undisplayed field), ADDR contains `const0_rtx'. If it is + not on the chain, ADDR contains `cc0_rtx'. ADDR is customarily + accessed with the macro `CONST_DOUBLE_MEM' and the chain field via + `CONST_DOUBLE_CHAIN'. + + If M is `VOIDmode', the bits of the value are stored in I0 and I1. + I0 is customarily accessed with the macro `CONST_DOUBLE_LOW' and + I1 with `CONST_DOUBLE_HIGH'. + + If the constant is floating point (regardless of its precision), + then the number of integers used to store the value depends on the + size of `REAL_VALUE_TYPE' (*note Cross-compilation::.). The + integers represent a floating point number, but not precisely in + the target machine's or host machine's floating point format. To + convert them to the precise bit pattern used by the target + machine, use the macro `REAL_VALUE_TO_TARGET_DOUBLE' and friends + (*note Data Output::.). + + The macro `CONST0_RTX (MODE)' refers to an expression with value 0 + in mode MODE. If mode MODE is of mode class `MODE_INT', it + returns `const0_rtx'. Otherwise, it returns a `CONST_DOUBLE' + expression in mode MODE. Similarly, the macro `CONST1_RTX (MODE)' + refers to an expression with value 1 in mode MODE and similarly + for `CONST2_RTX'. + +`(const_string STR)' + Represents a constant string with value STR. Currently this is + used only for insn attributes (*note Insn Attributes::.) since + constant strings in C are placed in memory. + +`(symbol_ref:MODE SYMBOL)' + Represents the value of an assembler label for data. SYMBOL is a + string that describes the name of the assembler label. If it + starts with a `*', the label is the rest of SYMBOL not including + the `*'. Otherwise, the label is SYMBOL, usually prefixed with + `_'. + + The `symbol_ref' contains a mode, which is usually `Pmode'. + Usually that is the only mode for which a symbol is directly valid. + +`(label_ref LABEL)' + Represents the value of an assembler label for code. It contains + one operand, an expression, which must be a `code_label' that + appears in the instruction sequence to identify the place where + the label should go. + + The reason for using a distinct expression type for code label + references is so that jump optimization can distinguish them. + +`(const:M EXP)' + Represents a constant that is the result of an assembly-time + arithmetic computation. The operand, EXP, is an expression that + contains only constants (`const_int', `symbol_ref' and `label_ref' + expressions) combined with `plus' and `minus'. However, not all + combinations are valid, since the assembler cannot do arbitrary + arithmetic on relocatable symbols. + + M should be `Pmode'. + +`(high:M EXP)' + Represents the high-order bits of EXP, usually a `symbol_ref'. + The number of bits is machine-dependent and is normally the number + of bits specified in an instruction that initializes the high + order bits of a register. It is used with `lo_sum' to represent + the typical two-instruction sequence used in RISC machines to + reference a global memory location. - The compiler assumes that certain kinds of RTL expressions are -unique; there do not exist two distinct objects representing the same -value. In other cases, it makes an opposite assumption: that no RTL -expression object of a certain kind appears in more than one place in -the containing structure. - - These assumptions refer to a single function; except for the RTL -objects that describe global variables and external functions, and a -few standard objects such as small integer constants, no RTL objects -are common to two functions. - - * Each pseudo-register has only a single `reg' object to represent - it, and therefore only a single machine mode. - - * For any symbolic label, there is only one `symbol_ref' object - referring to it. - - * There is only one `const_int' expression with value 0, only one - with value 1, and only one with value -1. Some other integer - values are also stored uniquely. - - * There is only one `pc' expression. - - * There is only one `cc0' expression. - - * There is only one `const_double' expression with value 0 for each - floating point mode. Likewise for values 1 and 2. - - * No `label_ref' or `scratch' appears in more than one place in the - RTL structure; in other words, it is safe to do a tree-walk of all - the insns in the function and assume that each time a `label_ref' - or `scratch' is seen it is distinct from all others that are seen. - - * Only one `mem' object is normally created for each static variable - or stack slot, so these objects are frequently shared in all the - places they appear. However, separate but equal objects for these - variables are occasionally made. - - * When a single `asm' statement has multiple output operands, a - distinct `asm_operands' expression is made for each output operand. - However, these all share the vector which contains the sequence of - input operands. This sharing is used later on to test whether two - `asm_operands' expressions come from the same statement, so all - optimizations must carefully preserve the sharing if they copy the - vector at all. - - * No RTL object appears in more than one place in the RTL structure - except as described above. Many passes of the compiler rely on - this by assuming that they can modify RTL objects in place without - unwanted side-effects on other insns. - - * During initial RTL generation, shared structure is freely - introduced. After all the RTL for a function has been generated, - all shared structure is copied by `unshare_all_rtl' in - `emit-rtl.c', after which the above rules are guaranteed to be - followed. - - * During the combiner pass, shared structure within an insn can exist - temporarily. However, the shared structure is copied before the - combiner is finished with the insn. This is done by calling - `copy_rtx_if_shared', which is a subroutine of `unshare_all_rtl'. + M should be `Pmode'.  -File: gcc.info, Node: Reading RTL, Prev: Sharing, Up: RTL +File: gcc.info, Node: Regs and Memory, Next: Arithmetic, Prev: Constants, Up: RTL -Reading RTL -=========== +Registers and Memory +==================== - To read an RTL object from a file, call `read_rtx'. It takes one -argument, a stdio stream, and returns a single RTL object. + Here are the RTL expression types for describing access to machine +registers and to main memory. - Reading RTL from a file is very slow. This is no currently not a -problem because reading RTL occurs only as part of building the -compiler. - - People frequently have the idea of using RTL stored as text in a -file as an interface between a language front end and the bulk of GNU -CC. This idea is not feasible. - - GNU CC was designed to use RTL internally only. Correct RTL for a -given program is very dependent on the particular target machine. And -the RTL does not contain all the information about the program. - - The proper way to interface GNU CC to a new language front end is -with the "tree" data structure. There is no manual for this data -structure, but it is described in the files `tree.h' and `tree.def'. +`(reg:M N)' + For small values of the integer N (those that are less than + `FIRST_PSEUDO_REGISTER'), this stands for a reference to machine + register number N: a "hard register". For larger values of N, it + stands for a temporary value or "pseudo register". The compiler's + strategy is to generate code assuming an unlimited number of such + pseudo registers, and later convert them into hard registers or + into memory references. + + M is the machine mode of the reference. It is necessary because + machines can generally refer to each register in more than one + mode. For example, a register may contain a full word but there + may be instructions to refer to it as a half word or as a single + byte, as well as instructions to refer to it as a floating point + number of various precisions. + + Even for a register that the machine can access in only one mode, + the mode must always be specified. + + The symbol `FIRST_PSEUDO_REGISTER' is defined by the machine + description, since the number of hard registers on the machine is + an invariant characteristic of the machine. Note, however, that + not all of the machine registers must be general registers. All + the machine registers that can be used for storage of data are + given hard register numbers, even those that can be used only in + certain instructions or can hold only certain types of data. + + A hard register may be accessed in various modes throughout one + function, but each pseudo register is given a natural mode and is + accessed only in that mode. When it is necessary to describe an + access to a pseudo register using a nonnatural mode, a `subreg' + expression is used. + + A `reg' expression with a machine mode that specifies more than + one word of data may actually stand for several consecutive + registers. If in addition the register number specifies a + hardware register, then it actually represents several consecutive + hardware registers starting with the specified one. + + Each pseudo register number used in a function's RTL code is + represented by a unique `reg' expression. + + Some pseudo register numbers, those within the range of + `FIRST_VIRTUAL_REGISTER' to `LAST_VIRTUAL_REGISTER' only appear + during the RTL generation phase and are eliminated before the + optimization phases. These represent locations in the stack frame + that cannot be determined until RTL generation for the function + has been completed. The following virtual register numbers are + defined: + + `VIRTUAL_INCOMING_ARGS_REGNUM' + This points to the first word of the incoming arguments + passed on the stack. Normally these arguments are placed + there by the caller, but the callee may have pushed some + arguments that were previously passed in registers. + + When RTL generation is complete, this virtual register is + replaced by the sum of the register given by + `ARG_POINTER_REGNUM' and the value of `FIRST_PARM_OFFSET'. + + `VIRTUAL_STACK_VARS_REGNUM' + If `FRAME_GROWS_DOWNWARD' is defined, this points to + immediately above the first variable on the stack. + Otherwise, it points to the first variable on the stack. + + `VIRTUAL_STACK_VARS_REGNUM' is replaced with the sum of the + register given by `FRAME_POINTER_REGNUM' and the value + `STARTING_FRAME_OFFSET'. + + `VIRTUAL_STACK_DYNAMIC_REGNUM' + This points to the location of dynamically allocated memory + on the stack immediately after the stack pointer has been + adjusted by the amount of memory desired. + + This virtual register is replaced by the sum of the register + given by `STACK_POINTER_REGNUM' and the value + `STACK_DYNAMIC_OFFSET'. + + `VIRTUAL_OUTGOING_ARGS_REGNUM' + This points to the location in the stack at which outgoing + arguments should be written when the stack is pre-pushed + (arguments pushed using push insns should always use + `STACK_POINTER_REGNUM'). + + This virtual register is replaced by the sum of the register + given by `STACK_POINTER_REGNUM' and the value + `STACK_POINTER_OFFSET'. + +`(subreg:M REG WORDNUM)' + `subreg' expressions are used to refer to a register in a machine + mode other than its natural one, or to refer to one register of a + multi-word `reg' that actually refers to several registers. + + Each pseudo-register has a natural mode. If it is necessary to + operate on it in a different mode--for example, to perform a + fullword move instruction on a pseudo-register that contains a + single byte--the pseudo-register must be enclosed in a `subreg'. + In such a case, WORDNUM is zero. + + Usually M is at least as narrow as the mode of REG, in which case + it is restricting consideration to only the bits of REG that are + in M. + + Sometimes M is wider than the mode of REG. These `subreg' + expressions are often called "paradoxical". They are used in + cases where we want to refer to an object in a wider mode but do + not care what value the additional bits have. The reload pass + ensures that paradoxical references are only made to hard + registers. + + The other use of `subreg' is to extract the individual registers of + a multi-register value. Machine modes such as `DImode' and + `TImode' can indicate values longer than a word, values which + usually require two or more consecutive registers. To access one + of the registers, use a `subreg' with mode `SImode' and a WORDNUM + that says which register. + + Storing in a non-paradoxical `subreg' has undefined results for + bits belonging to the same word as the `subreg'. This laxity makes + it easier to generate efficient code for such instructions. To + represent an instruction that preserves all the bits outside of + those in the `subreg', use `strict_low_part' around the `subreg'. + + The compilation parameter `WORDS_BIG_ENDIAN', if set to 1, says + that word number zero is the most significant part; otherwise, it + is the least significant part. + + Between the combiner pass and the reload pass, it is possible to + have a paradoxical `subreg' which contains a `mem' instead of a + `reg' as its first operand. After the reload pass, it is also + possible to have a non-paradoxical `subreg' which contains a + `mem'; this usually occurs when the `mem' is a stack slot which + replaced a pseudo register. + + Note that it is not valid to access a `DFmode' value in `SFmode' + using a `subreg'. On some machines the most significant part of a + `DFmode' value does not have the same format as a single-precision + floating value. + + It is also not valid to access a single word of a multi-word value + in a hard register when less registers can hold the value than + would be expected from its size. For example, some 32-bit + machines have floating-point registers that can hold an entire + `DFmode' value. If register 10 were such a register `(subreg:SI + (reg:DF 10) 1)' would be invalid because there is no way to + convert that reference to a single machine register. The reload + pass prevents `subreg' expressions such as these from being formed. + + The first operand of a `subreg' expression is customarily accessed + with the `SUBREG_REG' macro and the second operand is customarily + accessed with the `SUBREG_WORD' macro. + +`(scratch:M)' + This represents a scratch register that will be required for the + execution of a single instruction and not used subsequently. It is + converted into a `reg' by either the local register allocator or + the reload pass. + + `scratch' is usually present inside a `clobber' operation (*note + Side Effects::.). + +`(cc0)' + This refers to the machine's condition code register. It has no + operands and may not have a machine mode. There are two ways to + use it: + + * To stand for a complete set of condition code flags. This is + best on most machines, where each comparison sets the entire + series of flags. + + With this technique, `(cc0)' may be validly used in only two + contexts: as the destination of an assignment (in test and + compare instructions) and in comparison operators comparing + against zero (`const_int' with value zero; that is to say, + `const0_rtx'). + + * To stand for a single flag that is the result of a single + condition. This is useful on machines that have only a + single flag bit, and in which comparison instructions must + specify the condition to test. + + With this technique, `(cc0)' may be validly used in only two + contexts: as the destination of an assignment (in test and + compare instructions) where the source is a comparison + operator, and as the first operand of `if_then_else' (in a + conditional branch). + + There is only one expression object of code `cc0'; it is the value + of the variable `cc0_rtx'. Any attempt to create an expression of + code `cc0' will return `cc0_rtx'. + + Instructions can set the condition code implicitly. On many + machines, nearly all instructions set the condition code based on + the value that they compute or store. It is not necessary to + record these actions explicitly in the RTL because the machine + description includes a prescription for recognizing the + instructions that do so (by means of the macro + `NOTICE_UPDATE_CC'). *Note Condition Code::. Only instructions + whose sole purpose is to set the condition code, and instructions + that use the condition code, need mention `(cc0)'. + + On some machines, the condition code register is given a register + number and a `reg' is used instead of `(cc0)'. This is usually the + preferable approach if only a small subset of instructions modify + the condition code. Other machines store condition codes in + general registers; in such cases a pseudo register should be used. + + Some machines, such as the Sparc and RS/6000, have two sets of + arithmetic instructions, one that sets and one that does not set + the condition code. This is best handled by normally generating + the instruction that does not set the condition code, and making a + pattern that both performs the arithmetic and sets the condition + code register (which would not be `(cc0)' in this case). For + examples, search for `addcc' and `andcc' in `sparc.md'. + +`(pc)' + This represents the machine's program counter. It has no operands + and may not have a machine mode. `(pc)' may be validly used only + in certain specific contexts in jump instructions. + + There is only one expression object of code `pc'; it is the value + of the variable `pc_rtx'. Any attempt to create an expression of + code `pc' will return `pc_rtx'. + + All instructions that do not jump alter the program counter + implicitly by incrementing it, but there is no need to mention + this in the RTL. + +`(mem:M ADDR)' + This RTX represents a reference to main memory at an address + represented by the expression ADDR. M specifies how large a unit + of memory is accessed.  -File: gcc.info, Node: Machine Desc, Next: Target Macros, Prev: RTL, Up: Top +File: gcc.info, Node: Arithmetic, Next: Comparisons, Prev: Regs and Memory, Up: RTL -Machine Descriptions -******************** +RTL Expressions for Arithmetic +============================== - A machine description has two parts: a file of instruction patterns -(`.md' file) and a C header file of macro definitions. - - The `.md' file for a target machine contains a pattern for each -instruction that the target machine supports (or at least each -instruction that is worth telling the compiler about). It may also -contain comments. A semicolon causes the rest of the line to be a -comment, unless the semicolon is inside a quoted string. - - See the next chapter for information on the C header file. - -* Menu: - -* Patterns:: How to write instruction patterns. -* Example:: An explained example of a `define_insn' pattern. -* RTL Template:: The RTL template defines what insns match a pattern. -* Output Template:: The output template says how to make assembler code - from such an insn. -* Output Statement:: For more generality, write C code to output - the assembler code. -* Constraints:: When not all operands are general operands. -* Standard Names:: Names mark patterns to use for code generation. -* Pattern Ordering:: When the order of patterns makes a difference. -* Dependent Patterns:: Having one pattern may make you need another. -* Jump Patterns:: Special considerations for patterns for jump insns. -* Insn Canonicalizations::Canonicalization of Instructions -* Peephole Definitions::Defining machine-specific peephole optimizations. -* Expander Definitions::Generating a sequence of several RTL insns - for a standard operation. -* Insn Splitting:: Splitting Instructions into Multiple Instructions -* Insn Attributes:: Specifying the value of attributes for generated insns. + Unless otherwise specified, all the operands of arithmetic +expressions must be valid for mode M. An operand is valid for mode M +if it has mode M, or if it is a `const_int' or `const_double' and M is +a mode of class `MODE_INT'. + + For commutative binary operations, constants should be placed in the +second operand. + +`(plus:M X Y)' + Represents the sum of the values represented by X and Y carried + out in machine mode M. + +`(lo_sum:M X Y)' + Like `plus', except that it represents that sum of X and the + low-order bits of Y. The number of low order bits is + machine-dependent but is normally the number of bits in a `Pmode' + item minus the number of bits set by the `high' code (*note + Constants::.). + + M should be `Pmode'. + +`(minus:M X Y)' + Like `plus' but represents subtraction. + +`(compare:M X Y)' + Represents the result of subtracting Y from X for purposes of + comparison. The result is computed without overflow, as if with + infinite precision. + + Of course, machines can't really subtract with infinite precision. + However, they can pretend to do so when only the sign of the + result will be used, which is the case when the result is stored + in the condition code. And that is the only way this kind of + expression may validly be used: as a value to be stored in the + condition codes. + + The mode M is not related to the modes of X and Y, but instead is + the mode of the condition code value. If `(cc0)' is used, it is + `VOIDmode'. Otherwise it is some mode in class `MODE_CC', often + `CCmode'. *Note Condition Code::. + + Normally, X and Y must have the same mode. Otherwise, `compare' + is valid only if the mode of X is in class `MODE_INT' and Y is a + `const_int' or `const_double' with mode `VOIDmode'. The mode of X + determines what mode the comparison is to be done in; thus it must + not be `VOIDmode'. + + If one of the operands is a constant, it should be placed in the + second operand and the comparison code adjusted as appropriate. + + A `compare' specifying two `VOIDmode' constants is not valid since + there is no way to know in what mode the comparison is to be + performed; the comparison must either be folded during the + compilation or the first operand must be loaded into a register + while its mode is still known. + +`(neg:M X)' + Represents the negation (subtraction from zero) of the value + represented by X, carried out in mode M. + +`(mult:M X Y)' + Represents the signed product of the values represented by X and Y + carried out in machine mode M. + + Some machines support a multiplication that generates a product + wider than the operands. Write the pattern for this as + + (mult:M (sign_extend:M X) (sign_extend:M Y)) + + where M is wider than the modes of X and Y, which need not be the + same. + + Write patterns for unsigned widening multiplication similarly using + `zero_extend'. + +`(div:M X Y)' + Represents the quotient in signed division of X by Y, carried out + in machine mode M. If M is a floating point mode, it represents + the exact quotient; otherwise, the integerized quotient. + + Some machines have division instructions in which the operands and + quotient widths are not all the same; you should represent such + instructions using `truncate' and `sign_extend' as in, + + (truncate:M1 (div:M2 X (sign_extend:M2 Y))) + +`(udiv:M X Y)' + Like `div' but represents unsigned division. + +`(mod:M X Y)' +`(umod:M X Y)' + Like `div' and `udiv' but represent the remainder instead of the + quotient. + +`(smin:M X Y)' +`(smax:M X Y)' + Represents the smaller (for `smin') or larger (for `smax') of X + and Y, interpreted as signed integers in mode M. + +`(umin:M X Y)' +`(umax:M X Y)' + Like `smin' and `smax', but the values are interpreted as unsigned + integers. + +`(not:M X)' + Represents the bitwise complement of the value represented by X, + carried out in mode M, which must be a fixed-point machine mode. + +`(and:M X Y)' + Represents the bitwise logical-and of the values represented by X + and Y, carried out in machine mode M, which must be a fixed-point + machine mode. + +`(ior:M X Y)' + Represents the bitwise inclusive-or of the values represented by X + and Y, carried out in machine mode M, which must be a fixed-point + mode. + +`(xor:M X Y)' + Represents the bitwise exclusive-or of the values represented by X + and Y, carried out in machine mode M, which must be a fixed-point + mode. + +`(ashift:M X C)' + Represents the result of arithmetically shifting X left by C + places. X have mode M, a fixed-point machine mode. C be a + fixed-point mode or be a constant with mode `VOIDmode'; which mode + is determined by the mode called for in the machine description + entry for the left-shift instruction. For example, on the Vax, + the mode of C is `QImode' regardless of M. + +`(lshiftrt:M X C)' +`(ashiftrt:M X C)' + Like `ashift' but for right shift. Unlike the case for left shift, + these two operations are distinct. + +`(rotate:M X C)' +`(rotatert:M X C)' + Similar but represent left and right rotate. If C is a constant, + use `rotate'. + +`(abs:M X)' + Represents the absolute value of X, computed in mode M. + +`(sqrt:M X)' + Represents the square root of X, computed in mode M. Most often M + will be a floating point mode. + +`(ffs:M X)' + Represents one plus the index of the least significant 1-bit in X, + represented as an integer of mode M. (The value is zero if X is + zero.) The mode of X need not be M; depending on the target + machine, various mode combinations may be valid.  -File: gcc.info, Node: Patterns, Next: Example, Up: Machine Desc - -Everything about Instruction Patterns -===================================== +File: gcc.info, Node: Comparisons, Next: Bit Fields, Prev: Arithmetic, Up: RTL - Each instruction pattern contains an incomplete RTL expression, with -pieces to be filled in later, operand constraints that restrict how the -pieces can be filled in, and an output pattern or C code to generate -the assembler output, all wrapped up in a `define_insn' expression. - - A `define_insn' is an RTL expression containing four or five -operands: - - 1. An optional name. The presence of a name indicate that this - instruction pattern can perform a certain standard job for the - RTL-generation pass of the compiler. This pass knows certain - names and will use the instruction patterns with those names, if - the names are defined in the machine description. - - The absence of a name is indicated by writing an empty string - where the name should go. Nameless instruction patterns are never - used for generating RTL code, but they may permit several simpler - insns to be combined later on. - - Names that are not thus known and used in RTL-generation have no - effect; they are equivalent to no name at all. - - 2. The "RTL template" (*note RTL Template::.) is a vector of - incomplete RTL expressions which show what the instruction should - look like. It is incomplete because it may contain - `match_operand', `match_operator', and `match_dup' expressions - that stand for operands of the instruction. - - If the vector has only one element, that element is the template - for the instruction pattern. If the vector has multiple elements, - then the instruction pattern is a `parallel' expression containing - the elements described. - - 3. A condition. This is a string which contains a C expression that - is the final test to decide whether an insn body matches this - pattern. - - For a named pattern, the condition (if present) may not depend on - the data in the insn being matched, but only the - target-machine-type flags. The compiler needs to test these - conditions during initialization in order to learn exactly which - named instructions are available in a particular run. - - For nameless patterns, the condition is applied only when matching - an individual insn, and only after the insn has matched the - pattern's recognition template. The insn's operands may be found - in the vector `operands'. - - 4. The "output template": a string that says how to output matching - insns as assembler code. `%' in this string specifies where to - substitute the value of an operand. *Note Output Template::. +Comparison Operations +===================== - When simple substitution isn't general enough, you can specify a - piece of C code to compute the output. *Note Output Statement::. + Comparison operators test a relation on two operands and are +considered to represent a machine-dependent nonzero value described by, +but not necessarily equal to, `STORE_FLAG_VALUE' (*note Misc::.) if the +relation holds, or zero if it does not. The mode of the comparison +operation is independent of the mode of the data being compared. If +the comparison operation is being tested (e.g., the first operand of an +`if_then_else'), the mode must be `VOIDmode'. If the comparison +operation is producing data to be stored in some variable, the mode +must be in class `MODE_INT'. All comparison operations producing data +must use the same mode, which is machine-specific. + + There are two ways that comparison operations may be used. The +comparison operators may be used to compare the condition codes `(cc0)' +against zero, as in `(eq (cc0) (const_int 0))'. Such a construct +actually refers to the result of the preceding instruction in which the +condition codes were set. The instructing setting the condition code +must be adjacent to the instruction using the condition code; only +`note' insns may separate them. + + Alternatively, a comparison operation may directly compare two data +objects. The mode of the comparison is determined by the operands; they +must both be valid for a common machine mode. A comparison with both +operands constant would be invalid as the machine mode could not be +deduced from it, but such a comparison should never exist in RTL due to +constant folding. + + In the example above, if `(cc0)' were last set to `(compare X Y)', +the comparison operation is identical to `(eq X Y)'. Usually only one +style of comparisons is supported on a particular machine, but the +combine pass will try to merge the operations to produce the `eq' shown +in case it exists in the context of the particular insn involved. + + Inequality comparisons come in two flavors, signed and unsigned. +Thus, there are distinct expression codes `gt' and `gtu' for signed and +unsigned greater-than. These can produce different results for the same +pair of integer values: for example, 1 is signed greater-than -1 but not +unsigned greater-than, because -1 when regarded as unsigned is actually +`0xffffffff' which is greater than 1. + + The signed comparisons are also used for floating point values. +Floating point comparisons are distinguished by the machine modes of +the operands. + +`(eq:M X Y)' + 1 if the values represented by X and Y are equal, otherwise 0. + +`(ne:M X Y)' + 1 if the values represented by X and Y are not equal, otherwise 0. + +`(gt:M X Y)' + 1 if the X is greater than Y. If they are fixed-point, the + comparison is done in a signed sense. + +`(gtu:M X Y)' + Like `gt' but does unsigned comparison, on fixed-point numbers + only. + +`(lt:M X Y)' +`(ltu:M X Y)' + Like `gt' and `gtu' but test for "less than". + +`(ge:M X Y)' +`(geu:M X Y)' + Like `gt' and `gtu' but test for "greater than or equal". + +`(le:M X Y)' +`(leu:M X Y)' + Like `gt' and `gtu' but test for "less than or equal". + +`(if_then_else COND THEN ELSE)' + This is not a comparison operation but is listed here because it is + always used in conjunction with a comparison operation. To be + precise, COND is a comparison expression. This expression + represents a choice, according to COND, between the value + represented by THEN and the one represented by ELSE. + + On most machines, `if_then_else' expressions are valid only to + express conditional jumps. + +`(cond [TEST1 VALUE1 TEST2 VALUE2 ...] DEFAULT)' + Similar to `if_then_else', but more general. Each of TEST1, + TEST2, ... is performed in turn. The result of this expression is + the VALUE corresponding to the first non-zero test, or DEFAULT if + none of the tests are non-zero expressions. - 5. Optionally, a vector containing the values of attributes for insns - matching this pattern. *Note Insn Attributes::. + This is currently not valid for instruction patterns and is + supported only for insn attributes. *Note Insn Attributes::.  -File: gcc.info, Node: Example, Next: RTL Template, Prev: Patterns, Up: Machine Desc +File: gcc.info, Node: Bit Fields, Next: Conversions, Prev: Comparisons, Up: RTL + +Bit Fields +========== -Example of `define_insn' -======================== + Special expression codes exist to represent bitfield instructions. +These types of expressions are lvalues in RTL; they may appear on the +left side of an assignment, indicating insertion of a value into the +specified bit field. + +`(sign_extract:M LOC SIZE POS)' + This represents a reference to a sign-extended bit field contained + or starting in LOC (a memory or register reference). The bit field + is SIZE bits wide and starts at bit POS. The compilation option + `BITS_BIG_ENDIAN' says which end of the memory unit POS counts + from. + + If LOC is in memory, its mode must be a single-byte integer mode. + If LOC is in a register, the mode to use is specified by the + operand of the `insv' or `extv' pattern (*note Standard Names::.) + and is usually a full-word integer mode. + + The mode of POS is machine-specific and is also specified in the + `insv' or `extv' pattern. + + The mode M is the same as the mode that would be used for LOC if + it were a register. + +`(zero_extract:M LOC SIZE POS)' + Like `sign_extract' but refers to an unsigned or zero-extended bit + field. The same sequence of bits are extracted, but they are + filled to an entire word with zeros instead of by sign-extension. - Here is an actual example of an instruction pattern, for the -68000/68020. + +File: gcc.info, Node: Conversions, Next: RTL Declarations, Prev: Bit Fields, Up: RTL - (define_insn "tstsi" - [(set (cc0) - (match_operand:SI 0 "general_operand" "rm"))] - "" - "* - { if (TARGET_68020 || ! ADDRESS_REG_P (operands[0])) - return \"tstl %0\"; - return \"cmpl #0,%0\"; }") - - This is an instruction that sets the condition codes based on the -value of a general operand. It has no condition, so any insn whose RTL -description has the form shown may be handled according to this -pattern. The name `tstsi' means "test a `SImode' value" and tells the -RTL generation pass that, when it is necessary to test such a value, an -insn to do so can be constructed using this pattern. - - The output control string is a piece of C code which chooses which -output template to return based on the kind of operand and the specific -type of CPU for which code is being generated. +Conversions +=========== - `"rm"' is an operand constraint. Its meaning is explained below. + All conversions between machine modes must be represented by +explicit conversion operations. For example, an expression which is +the sum of a byte and a full word cannot be written as `(plus:SI +(reg:QI 34) (reg:SI 80))' because the `plus' operation requires two +operands of the same machine mode. Therefore, the byte-sized operand +is enclosed in a conversion operation, as in + + (plus:SI (sign_extend:SI (reg:QI 34)) (reg:SI 80)) + + The conversion operation is not a mere placeholder, because there +may be more than one way of converting from a given starting mode to +the desired final mode. The conversion operation code says how to do +it. + + For all conversion operations, X must not be `VOIDmode' because the +mode in which to do the conversion would not be known. The conversion +must either be done at compile-time or X must be placed into a register. + +`(sign_extend:M X)' + Represents the result of sign-extending the value X to machine + mode M. M must be a fixed-point mode and X a fixed-point value of + a mode narrower than M. + +`(zero_extend:M X)' + Represents the result of zero-extending the value X to machine + mode M. M must be a fixed-point mode and X a fixed-point value of + a mode narrower than M. + +`(float_extend:M X)' + Represents the result of extending the value X to machine mode M. + m must be a floating point mode and X a floating point value of a + mode narrower than M. + +`(truncate:M X)' + Represents the result of truncating the value X to machine mode M. + M must be a fixed-point mode and X a fixed-point value of a mode + wider than M. + +`(float_truncate:M X)' + Represents the result of truncating the value X to machine mode M. + M must be a floating point mode and X a floating point value of a + mode wider than M. + +`(float:M X)' + Represents the result of converting fixed point value X, regarded + as signed, to floating point mode M. + +`(unsigned_float:M X)' + Represents the result of converting fixed point value X, regarded + as unsigned, to floating point mode M. + +`(fix:M X)' + When M is a fixed point mode, represents the result of converting + floating point value X to mode M, regarded as signed. How + rounding is done is not specified, so this operation may be used + validly in compiling C code only for integer-valued operands. + +`(unsigned_fix:M X)' + Represents the result of converting floating point value X to + fixed point mode M, regarded as unsigned. How rounding is done is + not specified. + +`(fix:M X)' + When M is a floating point mode, represents the result of + converting floating point value X (valid for mode M) to an + integer, still represented in floating point mode M, by rounding + towards zero.  -File: gcc.info, Node: RTL Template, Next: Output Template, Prev: Example, Up: Machine Desc +File: gcc.info, Node: RTL Declarations, Next: Side Effects, Prev: Conversions, Up: RTL -RTL Template +Declarations ============ - The RTL template is used to define which insns match the particular -pattern and how to find their operands. For named patterns, the RTL -template also says how to construct an insn from specified operands. - - Construction involves substituting specified operands into a copy of -the template. Matching involves determining the values that serve as -the operands in the insn being matched. Both of these activities are -controlled by special expression types that direct matching and -substitution of the operands. - -`(match_operand:M N PREDICATE CONSTRAINT)' - This expression is a placeholder for operand number N of the insn. - When constructing an insn, operand number N will be substituted - at this point. When matching an insn, whatever appears at this - position in the insn will be taken as operand number N; but it - must satisfy PREDICATE or this instruction pattern will not match - at all. - - Operand numbers must be chosen consecutively counting from zero in - each instruction pattern. There may be only one `match_operand' - expression in the pattern for each operand number. Usually - operands are numbered in the order of appearance in `match_operand' - expressions. - - PREDICATE is a string that is the name of a C function that - accepts two arguments, an expression and a machine mode. During - matching, the function will be called with the putative operand as - the expression and M as the mode argument (if M is not specified, - `VOIDmode' will be used, which normally causes PREDICATE to accept - any mode). If it returns zero, this instruction pattern fails to - match. PREDICATE may be an empty string; then it means no test is - to be done on the operand, so anything which occurs in this - position is valid. - - Most of the time, PREDICATE will reject modes other than M--but - not always. For example, the predicate `address_operand' uses M - as the mode of memory ref that the address should be valid for. - Many predicates accept `const_int' nodes even though their mode is - `VOIDmode'. - - CONSTRAINT controls reloading and the choice of the best register - class to use for a value, as explained later (*note - Constraints::.). - - People are often unclear on the difference between the constraint - and the predicate. The predicate helps decide whether a given - insn matches the pattern. The constraint plays no role in this - decision; instead, it controls various decisions in the case of an - insn which does match. - - On CISC machines, the most common PREDICATE is - `"general_operand"'. This function checks that the putative - operand is either a constant, a register or a memory reference, - and that it is valid for mode M. - - For an operand that must be a register, PREDICATE should be - `"register_operand"'. Using `"general_operand"' would be valid, - since the reload pass would copy any non-register operands through - registers, but this would make GNU CC do extra work, it would - prevent invariant operands (such as constant) from being removed - from loops, and it would prevent the register allocator from doing - the best possible job. On RISC machines, it is usually most - efficient to allow PREDICATE to accept only objects that the - constraints allow. - - For an operand that must be a constant, you must be sure to either - use `"immediate_operand"' for PREDICATE, or make the instruction - pattern's extra condition require a constant, or both. You cannot - expect the constraints to do this work! If the constraints allow - only constants, but the predicate allows something else, the - compiler will crash when that case arises. - -`(match_scratch:M N CONSTRAINT)' - This expression is also a placeholder for operand number N and - indicates that operand must be a `scratch' or `reg' expression. - - When matching patterns, this is equivalent to - - (match_operand:M N "scratch_operand" PRED) - - but, when generating RTL, it produces a (`scratch':M) expression. - - If the last few expressions in a `parallel' are `clobber' - expressions whose operands are either a hard register or - `match_scratch', the combiner can add or delete them when - necessary. *Note Side Effects::. - -`(match_dup N)' - This expression is also a placeholder for operand number N. It is - used when the operand needs to appear more than once in the insn. - - In construction, `match_dup' acts just like `match_operand': the - operand is substituted into the insn being constructed. But in - matching, `match_dup' behaves differently. It assumes that operand - number N has already been determined by a `match_operand' - appearing earlier in the recognition template, and it matches only - an identical-looking expression. - -`(match_operator:M N PREDICATE [OPERANDS...])' - This pattern is a kind of placeholder for a variable RTL expression - code. - - When constructing an insn, it stands for an RTL expression whose - expression code is taken from that of operand N, and whose - operands are constructed from the patterns OPERANDS. - - When matching an expression, it matches an expression if the - function PREDICATE returns nonzero on that expression *and* the - patterns OPERANDS match the operands of the expression. - - Suppose that the function `commutative_operator' is defined as - follows, to match any expression whose operator is one of the - commutative arithmetic operators of RTL and whose mode is MODE: - - int - commutative_operator (x, mode) - rtx x; - enum machine_mode mode; - { - enum rtx_code code = GET_CODE (x); - if (GET_MODE (x) != mode) - return 0; - return (GET_RTX_CLASS (code) == 'c' - || code == EQ || code == NE); - } - - Then the following pattern will match any RTL expression consisting - of a commutative operator applied to two general operands: - - (match_operator:SI 3 "commutative_operator" - [(match_operand:SI 1 "general_operand" "g") - (match_operand:SI 2 "general_operand" "g")]) - - Here the vector `[OPERANDS...]' contains two patterns because the - expressions to be matched all contain two operands. - - When this pattern does match, the two operands of the commutative - operator are recorded as operands 1 and 2 of the insn. (This is - done by the two instances of `match_operand'.) Operand 3 of the - insn will be the entire commutative expression: use `GET_CODE - (operands[3])' to see which commutative operator was used. - - The machine mode M of `match_operator' works like that of - `match_operand': it is passed as the second argument to the - predicate function, and that function is solely responsible for - deciding whether the expression to be matched "has" that mode. - - When constructing an insn, argument 3 of the gen-function will - specify the operation (i.e. the expression code) for the - expression to be made. It should be an RTL expression, whose - expression code is copied into a new expression whose operands are - arguments 1 and 2 of the gen-function. The subexpressions of - argument 3 are not used; only its expression code matters. - - When `match_operator' is used in a pattern for matching an insn, - it usually best if the operand number of the `match_operator' is - higher than that of the actual operands of the insn. This improves - register allocation because the register allocator often looks at - operands 1 and 2 of insns to see if it can do register tying. - - There is no way to specify constraints in `match_operator'. The - operand of the insn which corresponds to the `match_operator' - never has any constraints because it is never reloaded as a whole. - However, if parts of its OPERANDS are matched by `match_operand' - patterns, those parts may have constraints of their own. - -`(match_op_dup:M N[OPERANDS...])' - Like `match_dup', except that it applies to operators instead of - operands. When constructing an insn, operand number N will be - substituted at this point. But in matching, `match_op_dup' behaves - differently. It assumes that operand number N has already been - determined by a `match_operator' appearing earlier in the - recognition template, and it matches only an identical-looking - expression. + Declaration expression codes do not represent arithmetic operations +but rather state assertions about their operands. -`(match_parallel N PREDICATE [SUBPAT...])' - This pattern is a placeholder for an insn that consists of a - `parallel' expression with a variable number of elements. This - expression should only appear at the top level of an insn pattern. - - When constructing an insn, operand number N will be substituted at - this point. When matching an insn, it matches if the body of the - insn is a `parallel' expression with at least as many elements as - the vector of SUBPAT expressions in the `match_parallel', if each - SUBPAT matches the corresponding element of the `parallel', *and* - the function PREDICATE returns nonzero on the `parallel' that is - the body of the insn. It is the responsibility of the predicate - to validate elements of the `parallel' beyond those listed in the - `match_parallel'. - - A typical use of `match_parallel' is to match load and store - multiple expressions, which can contain a variable number of - elements in a `parallel'. For example, - - (define_insn "" - [(match_parallel 0 "load_multiple_operation" - [(set (match_operand:SI 1 "gpc_reg_operand" "=r") - (match_operand:SI 2 "memory_operand" "m")) - (use (reg:SI 179)) - (clobber (reg:SI 179))])] - "" - "loadm 0,0,%1,%2") - - This example comes from `a29k.md'. The function - `load_multiple_operations' is defined in `a29k.c' and checks that - subsequent elements in the `parallel' are the same as the `set' in - the pattern, except that they are referencing subsequent registers - and memory locations. - - An insn that matches this pattern might look like: - - (parallel - [(set (reg:SI 20) (mem:SI (reg:SI 100))) - (use (reg:SI 179)) - (clobber (reg:SI 179)) - (set (reg:SI 21) - (mem:SI (plus:SI (reg:SI 100) - (const_int 4)))) - (set (reg:SI 22) - (mem:SI (plus:SI (reg:SI 100) - (const_int 8))))]) - -`(match_par_dup N [SUBPAT...])' - Like `match_op_dup', but for `match_parallel' instead of - `match_operator'. - -`(address (match_operand:M N "address_operand" ""))' - This complex of expressions is a placeholder for an operand number - N in a "load address" instruction: an operand which specifies a - memory location in the usual way, but for which the actual operand - value used is the address of the location, not the contents of the - location. - - `address' expressions never appear in RTL code, only in machine - descriptions. And they are used only in machine descriptions that - do not use the operand constraint feature. When operand - constraints are in use, the letter `p' in the constraint serves - this purpose. - - M is the machine mode of the *memory location being addressed*, - not the machine mode of the address itself. That mode is always - the same on a given target machine (it is `Pmode', which normally - is `SImode'), so there is no point in mentioning it; thus, no - machine mode is written in the `address' expression. If some day - support is added for machines in which addresses of different - kinds of objects appear differently or are used differently (such - as the PDP-10), different formats would perhaps need different - machine modes and these modes might be written in the `address' +`(strict_low_part (subreg:M (reg:N R) 0))' + This expression code is used in only one context: as the + destination operand of a `set' expression. In addition, the + operand of this expression must be a non-paradoxical `subreg' expression. - -File: gcc.info, Node: Output Template, Next: Output Statement, Prev: RTL Template, Up: Machine Desc - -Output Templates and Operand Substitution -========================================= - - The "output template" is a string which specifies how to output the -assembler code for an instruction pattern. Most of the template is a -fixed string which is output literally. The character `%' is used to -specify where to substitute an operand; it can also be used to identify -places where different variants of the assembler require different -syntax. - - In the simplest case, a `%' followed by a digit N says to output -operand N at that point in the string. - - `%' followed by a letter and a digit says to output an operand in an -alternate fashion. Four letters have standard, built-in meanings -described below. The machine description macro `PRINT_OPERAND' can -define additional letters with nonstandard meanings. - - `%cDIGIT' can be used to substitute an operand that is a constant -value without the syntax that normally indicates an immediate operand. - - `%nDIGIT' is like `%cDIGIT' except that the value of the constant is -negated before printing. - - `%aDIGIT' can be used to substitute an operand as if it were a -memory reference, with the actual operand treated as the address. This -may be useful when outputting a "load address" instruction, because -often the assembler syntax for such an instruction requires you to -write the operand as if it were a memory reference. - - `%lDIGIT' is used to substitute a `label_ref' into a jump -instruction. - - `%=' outputs a number which is unique to each instruction in the -entire compilation. This is useful for making local labels to be -referred to more than once in a single template that generates multiple -assembler instructions. - - `%' followed by a punctuation character specifies a substitution that -does not use an operand. Only one case is standard: `%%' outputs a `%' -into the assembler code. Other nonstandard cases can be defined in the -`PRINT_OPERAND' macro. You must also define which punctuation -characters are valid with the `PRINT_OPERAND_PUNCT_VALID_P' macro. - - The template may generate multiple assembler instructions. Write -the text for the instructions, with `\;' between them. - - When the RTL contains two operands which are required by constraint -to match each other, the output template must refer only to the -lower-numbered operand. Matching operands are not always identical, -and the rest of the compiler arranges to put the proper RTL expression -for printing into the lower-numbered operand. - - One use of nonstandard letters or punctuation following `%' is to -distinguish between different assembler languages for the same machine; -for example, Motorola syntax versus MIT syntax for the 68000. Motorola -syntax requires periods in most opcode names, while MIT syntax does -not. For example, the opcode `movel' in MIT syntax is `move.l' in -Motorola syntax. The same file of patterns is used for both kinds of -output syntax, but the character sequence `%.' is used in each place -where Motorola syntax wants a period. The `PRINT_OPERAND' macro for -Motorola syntax defines the sequence to output a period; the macro for -MIT syntax defines it to do nothing. - - As a special case, a template consisting of the single character `#' -instructs the compiler to first split the insn, and then output the -resulting instructions separately. This helps eliminate redundancy in -the output templates. If you have a `define_insn' that needs to emit -multiple assembler instructions, and there is an matching `define_split' -already defined, then you can simply use `#' as the output template -instead of writing an output template that emits the multiple assembler -instructions. - - If `ASSEMBLER_DIALECT' is defined, you can use -`{option0|option1|option2}' constructs in the templates. These -describe multiple variants of assembler language syntax. *Note -Instruction Output::. + The presence of `strict_low_part' says that the part of the + register which is meaningful in mode N, but is not part of mode M, + is not to be altered. Normally, an assignment to such a subreg is + allowed to have undefined effects on the rest of the register when + M is less than a word.