UNIT 1 Ca

P. R.
Engineering College
Department of CSE & IT
P.R.ENGINEERING COLLEGE
AICTE APPROVED AND AFFILIATED TO ANNA UNIVERSITY
VALLAM, THANJAVUR
DEPARTMENT OF CSE AND IT
LECTURE NOTES AND QUESTION BANK FOR
CS6303 COMPUTER ARCHITECTURE
PREPARED BY
N. Mohammed Haris
ASSISTANT PROFESSOR
DEPARTMENT OF IT
P. R. Engineering College
SYLABUS
CS6303 COMPUTER ARCHITECTURE
LTPC
3003
UNIT I OVERVIEW & INSTRUCTIONS 9

Eight ideas Components of a computer system Technology Performance Power wall
Uniprocessors to multiprocessors; Instructions operations and operands representing instructions
Logical operations control operations Addressing and addressing modes.
UNIT II ARITHMETIC OPERATIONS 7
ALU - Addition and subtraction Multiplication Division Floating Point operations Subword
parallelism.
UNIT III PROCESSOR AND CONTROL UNIT 11
Basic MIPS implementation Building datapath Control Implementation scheme Pipelining
Pipelined datapath and control Handling Data hazards & Control hazards Exceptions.
UNIT IV PARALLELISM 9
Instruction-level-parallelism Parallel processing challenges Flynn's classification Hardware
multithreading Multicore processors
36
UNIT V MEMORY AND I/O SYSTEMS 9
Memory hierarchy - Memory technologies Cache basics Measuring and improving cache
performance - Virtual memory, TLBs - Input/output system, programmed I/O, DMA and interrupts, I/O
processors.
TOTAL: 45 PERIODS
TEXT BOOK:
1. David A. Patterson and John L. Hennessey, Computer organization and design, Morgan
Kauffman / Elsevier, Fifth edition, 2014.
REFERENCES:
1. V.Carl Hamacher, Zvonko G. Varanesic and Safat G. Zaky, Computer Organisation,
VI th edition, Mc Graw-Hill Inc, 2012.
2. William Stallings Computer Organization and Architecture , Seventh Edition , Pearson
Education, 2006.
3. Vincent P. Heuring, Harry F. Jordan, Computer System Architecture, Second Edition,
Pearson Education, 2005.
4. Govindarajalu, Computer Architecture and Organization, Design Principles and Applications",
first edition, Tata McGraw Hill, New Delhi, 2005.
5. John P. Hayes, Computer Architecture and Organization, Third Edition, Tata Mc Graw Hill,
1998.
6. http://nptel.ac.in/.
2
UNIT I - OVERVIEW & INSTRUCTIONS

Computer Architecture:
The overall design or structure of a computer system, including the hardware and the
software required to run it, especially the internal structure of the microprocessor.
EIGHT GREAT IDEAS
1.1
Eight Great Ideas

_ Design for Moores Law
_ Use abstraction to simplify design
_ Make the common case fast
_ Performance via parallelism
_ Performance via pipelining
_ Performance via prediction
_ Hierarchy of memories
_ Dependability via redundancy
1.1.1
Design for Moores Law

The one constant for computer designers is rapid change, which is driven largely by
Moores Law.
It states that integrated circuit resources double every 1824 months.
As computer designs can take years, the resources available per chip can easily double
or quadruple between the start and finish of the project.
1.1.2 Use Abstraction to Simplify Design
A major productivity technique for hardware and software is to use abstractions.
To represent the design at different levels of representation; lower-levels and higher

levels
Lower level details are hidden to offer a simpler model at higher levels.
1.1.3 Make the Common Case Fast
Making the common case fast will tend to enhance performance better than optimizing
the rare case.
Ironically, the common case is often simpler than the rare case and hence is often
easier to enhance.
1.1.4
Performance via Parallelism

Computer architects have offered designs that get more performance by performing
operations in parallel.
1.1.5
Performance via Pipelining

3
1.1.6
A particular pattern of parallelism is pipelining.

Performance via Prediction
In some cases. it can be faster on average to guess and start working rather than wait
until you know for sure.
Assuming that the mechanism to recover from a misprediction is not too expensive and
your prediction is relatively accurate.
1.1.7
Hierarchy of Memories
The Computer Architects use a layered triangle icon to represent the memory hierarchy.
Figure 1.1 Memory hierarchy

The shape indicates speed, cost, and size (From Figure 1.1 )
(i) The closer to the top, the faster and more expensive per bit the memory;
(ii) The wider the base of the layer, the bigger the memory.
1.1.8
Dependability via Redundancy
Computers not only need to be fast; they need to be dependable.
Since any physical device can fail, we make systems dependable by including
redundant components that can take over when a failure occurs and to help detect
failures.
1.2 COMPONENTS OF A COMPUTER SYSTEM (Functional Units)

_ Same components for all kinds of computers
_ Desktop, server, embedded
_ Input/output includes
_ User-interface devices
_ Display, keyboard, mouse
_ Storage devices
_ Hard disk, CD/DVD, flash
_ Network adapters
_ For communicating with other computers
A modern computer consisting of hardware and software.
The hardware has five different types of functional units (figure: 1.2): Memory unit, Data
path unit (ALU), Control unit, input unit and output unit.
Input unit - A mechanism through which the computer is fed information. E.g. Keyboard,
Joystick, Trackball and Mouse
Mouse:
The original mouse was electromechanical and used a large ball that when rolled
across a surface would cause an x and y counter to be incremented.
The amount of increase in each counter told how far the mouse had been moved.
Nowadays, the electromechanical mouse is replaced by optical mouse.
The replacement of optical mouse, reduce cost and increase reliability.
It includes Light emitting diode (LED) to provide lighting, a tiny black and white
camera.
The LED is underneath the mouse, the camera takes 1500 sample pictures/second.
The sample pictures are sent to optical processor that compare the images and
determine how far the mouse has moved.
Central Processor Unit (CPU) - Also called processor. The active part of the computer,
which contains the datapath and control units.
Datapath unit - It performs the arithmetic operations.
Registers:
Registers are faster than memory.
Registers can hold variables and intermediate results.
Memory traffic is reduced, so program runs faster.
Control unit It fetches and analyses the instructions one-by-one and issue control
signals to all other units to perform various operations.
For a given instruction, the exact set of operations required is indicated by the control
signals. The results of instructions are stored in memory.
Figure 1.2: Components of computer system
Memory unit:
5
It stores the program and data.

There are 2 classes of memory.1.Primary, 2.Secondary.
Primary memory also called main memory. Memory used to hold programs while
they are running. E.g.DRAM (Dynamic Random access memory).
The memory inside the computer is volatile storage, such as DRAM, that retains
data only if it is receiving power.
Secondary memory (Nonvolatile storage) is a form of storage that retains data
even in the absence of a power source and that is used to store the programs
between runs.E.g.Magnetic disks.
Types:
1. Volatile Memory: Data lost when the power turns off and that is used to hold
data and program while they are running.
Eg: DRAM
2. Non-Volatile Memory: It retains data even in the absence of a power source and
that is store programs between runs.
Eg: Magnetic disk, Flash memory, Optical Disk
Magnetic Disk consist of a collection of platters, which rotate on a spindle
at 5400 revolution/min.
The metal platter are covered with magnetic recording material on both sides.
Optical Disk: Include both Compact Disk(CD) and Digital Video Disk(DVD)
Read-Only CD/DVD:
o Data is recorded in a spiral fashion, with individual bits being recorded
by burning small pit.
o The disk is read by shining a laser at the CD surface and determining
by examining the reflected light whether there is a pit or flat surface.
Rewritable CD/DVD:
Use different recording surface that as a crystal line, reflective
materials, pits are formed that are not reflective.
Erase CD/DVD:
The surface is heated and cooled slowly, allowing an annealing
process to restore the surface recording layer to its crystalline structure.
Output unit - A mechanism that conveys the result of a computation to a user, such as a
display, or to another computer. E.g.Printer,Floppy Disk, DVD
6
Monitor:
All laptops and desktop computers use Liquid Crystal Display (LCD) to get a thin,
low-power display.
A tiny transistor switch at each pixel to control current and make sharper images
The image is composed of a matrix of picture elements or pixels, which can be

represented as a matrix of bits called bitmap.
Depending on the size of the screen and the resolution, the display matrix ranges in
size from 640*480 t0 2560*1600.
A red, green, blue (RGB) associated with each dot on the display determines the
intensity of the three color components in the final image.
Networks
_ Communication, resource sharing, nonlocal access
_ Local area network (LAN): Ethernet
_ Wide area network (WAN): the Internet
_ Wireless network: WiFi, Bluetooth
1.3 TECHNOLOGIES FOR BUILDING PROCESSORS AND MEMORY
The technology shapes what computers will be able to do and how

quickly they will evolve.
Year
Technology
Relative performance/cost
1951
Vacuum tube
1965
Transistor
35
1975
Integrated circuit (IC)
900
1995
Very large scale IC (VLSI)
2,400,000
2013
Ultra large scale IC(ULSI)
250,000,000,000
A transistor is simply an on/off switch controlled by electricity.

7
The integrated circuit (IC) combined dozens to hundreds of transistors into a single
chip.
Very Large-Scale Integrated (VLSI) circuit-A device containing hundreds of

thousands to millions of transistors.
Manufacturing IC
Intel I7 Wafer
Integrated Circuit Cost

_ Nonlinear relation to area and defect rate
_ Wafer cost and area are fixed
_ Defect rate determined by manufacturing process
_ Die area determined by architecture and circuit design
8
Ultra Large-Scale Integrated

, CPU execution time, and so on.
Throughput - Also called bandwidth. It is the number of tasks completed per unit
time.
To maximize performance, we want to minimize response time or execution time for

some task.
1.4 Performance
Performance assessment among different computers is quite complicated to purchasers
and to designers. Therefore, understanding how to measure performance and the limitations of
performance measurements is important in selecting a computer.
1. Performance Evaluation
Computer performance evaluation is primarily based on throughput and response time.
The computer user is interested in reducing response timethe time between the start
and the completion of an eventalso referred to as execution time.
The manager of a large data processing center may be interested in increasing

throughputthe total amount of work done in a given time.
To maximize performance, we want to minimize response time or execution time for

some task.
The relation between performance and execution time for a computer X:

--------------- (1.4.1)
This means that for two computers X and Y, if the performance of X is greater than the
performance of Y, we have
--------------- (1.4.2)
--------------- (1.4.3)
--------------- (1.4.4)
That is, the execution time on Y is longer than that on X, if X is faster than Y.
--------------- (1.4.5)
If X is n times as fast as Y, then the execution time on Y is n times as long as it is on X:
er instruction.
9
Therefore, the number of clock cycles required for a program can be written as
----(1.4.9)
Clock cycles per instruction (CPI)-Average number of clock cycles per instruction for
a program or program fragment.
Table 1: The basic components of performance and its units of measure.
S.No
Components of Performance
Units of measure
CPU execution time for a program
Seconds for the program
Instruction Count
Instructions Executed for the Program
Clock cycles per instruction (CPI)
Average number of clock cycles per instruction
Clock cycle time
Seconds per clock cycle
The Classic CPU Performance Equation

Relative Performance Example
_ If computer A runs a program in 10 seconds and computer B runs the same program in 15
seconds, how much faster is A than B?
We know that A is n times faster than B if
Measuring Execution Time

_ Elapsed time
_ Total response time, including all aspects
_ Processing, I/O, OS overhead, idle time
_ Determines system performance
_ CPU time
_ Time spent processing a given job
_ Comprises user CPU time and system CPU time
_ Different programs are affected differently by CPU and system performance
10
_ How do we compare one computer/processor against another?

Performance Metrics
_ Purchasing perspective
_ Given a collection of machines, which has the
_ Best performance ?
_ Least cost ?
_ Best cost/performance?
_ Design perspective
_ Faced with design options, which has the
_ Best performance improvement ?
_ Least cost ?
_ Best cost/performance?
_ Both require
_ Basis for comparison
_ Metric for evaluation
_ Our goal in this class is to understand what factors in the architecture contribute to overall
system performance and the relative importance (and cost) of these factors
Performance Factors
_ CPU execution time (CPU time) time the CPU spends working on a task
_ Does not include time waiting for I/O or running other programs
_ We can improve performance by reducing either the length of the clock cycle or the number
of clock cycles required for a program
_ Hardware designers must often trade off clock rate against cycle count
Using the Performance Equation
_ Computers A and B implement the same Instruction Set Architecture (ISA). Computer A has
a clock cycle time of 250 ps and an effective CPI of 2.0 for some program and computer B has
a clock cycle time of 500 ps and an effective CPI of 1.2 for the same program. Which computer
is faster and by how much?
Each computer executes the same number of instructions, I, so
CPU timeA = I x 2.0 x 250 ps = 500 x I ps
CPU timeB = I x 1.2 x 500 ps = 600 x I ps
Clearly, A is faster by the ratio of execution times
11
Improving Performance Example

_ A program runs on computer A with a 2 GHz clock in 10 seconds. What clock rate must
computer B run at to run this program in 6 seconds? Unfortunately, to accomplish this,
computer B will require 1.2 times as many clock cycles as computer A to run the program.
Instruction Count and CPI
_ Instruction Count for a program

_ Determined by program, ISA, and compiler
_ Average cycles per instruction
_ Determined by CPU hardware
_ If different instructions have different CPI
_ Average CPI affected by instruction mix
Effective (Average) CPI
_ Computing the overall effective CPI is done by looking at the different types of instructions
and their individual cycle counts and averaging
_ Where ICi is the count (percentage) of the number of instructions of class i executed
_ CPIi is the (average) number of clock cycles per instruction for that instruction class
12
_ n is the number of instruction classes

_ The overall effective CPI varies by instruction mix a measure of the dynamic frequency of
instructions across one or many programs
THE Performance Equation
_ Our basic performance equation is then
CPU time = Instruction_count x CPI x clock_cycle
_ These equations separate the three key factors that affect performance
_ Can measure the CPU execution time by running the program
_ The clock rate is usually given
_ Can measure overall instruction count by using profilers/ simulators
_ CPI varies by instruction type and ISA implementation
Performance Summary
_ Performance depends on
_ Algorithm: affects IC, possibly CPI
_ Programming language: affects IC, CPI
_ Compiler: affects IC, CPI
_ Instruction set architecture: affects IC, CPI, Tc
1.5 POWER WALL
The problem in today technology is the SPEC performance of the hottest chip grew by
52% per year from 1986 to 2002 and then grew only 20% in the next three years.
The problem is now called the POWER WALL.
Goal: Drive clock rate up
Done by adding more transistors to a smaller chip
Increased the power dissipation of the CPU chip beyond the capacity of inexpensive
cooling techniques
Both clock rate and power increased rapidly for decades, and then flattened off recently.
13
The reason they grew together is that they are correlated, and the reason for their
recent slowing is that we have run into the practical power limit for cooling commodity
microprocessors.
The dominant technology for integrated circuits is called CMO S (complementary metal
oxide semiconductor).
For CMOS, the primary source of energy consumption is so-called dynamic energy.
Dynamic energy- energy that is consumed when transistors switch states from 0 to 1
and vice versa. It is depends on the capacitive loading of each transistor and the voltage
applied:
CMOS technology
Reduce power by reducing frequency
Hence, recent processors intended for laptop use all have the ability to adapt
frequency to reduce power consumption, simultaneously, of course, reducing
performance.
Thus, adequately evaluating the energy efficiency of a processor requires

examining its performance at maximum power, at an intermediate level that conserves
battery life, and at a level that maximizes battery life.
Two available clock rates: maximum and a reduced clock rate.

The best performance is obtained by running at maximum speed, the best battery
life by running always at the reduced rate, and the intermediate, performance-power
optimized level by switching dynamically between these two clock rates.
----- (1.5.1)
14
This equation is the energy of a pulse during the logic transition of 0 1 0 or 1 0

1. The energy of a single transition is then
----- (1.5.2)
The power required per transistor is just the product of energy of a transition and the
frequency of transitions:
----- (1.5.3)
Frequency switched it is a function of the clock rate.
The capacitive load per transistor is a function of both the number of transistors
connected to an output (called the fanout) and the technology, which determines the
capacitance of both wires and transistors.
1.6 UNIPROCESSORS TO MULTIPROCESSORS

The power wall has forced a dramatic change in the design of microprocessor.
Since 2022, the rate has slowed from a factor of 1.5/year to less than a factor of 1.2/year.
Rather than continuing to decrease the response time of a single program running on the
single-processor, as of 2006 all desktop and server companies are shipping with
multiprocessor per chip, where the benefit is often more on throughput than on response
time.
Product
Core per chip
AMD
Intel
IBM Power 6
Operation X4
Nahalem
Sun Ultra
Clock rate
2.5Ghz
~2.5Ghz
4.7Ghz
1.4Ghz
Microprocessor
120W
~100W
~100W
94 W
power
The multicore uses more than processor and require explicitly parallel hardware and
programming.
Parallel hardware means that the programmer must divide an application so that each
processor has roughly the same amount to do at the same time, and the overhead of
15
scheduling and coordination doesnt waste energy away the potential performance benefits
of parallelism.
In multiprocessing, the load is evenly between the parallel hardware to get the desired
speedup with proper scheduling the task to each hardware unit.
If one hardware module takes longer time than the other one, then other hardware module
should wait until all hardware modules were completed.
Thus, care must be taken to reduce communication and synchronization overhead.
The power limit has forced a dramatic change in the design of microprocessors.
There is a decrease in the response time of a single program running on the single
processor, as of 2006 a switch occurred are from uniprocessors to microprocessors with
multiple processors per chip.
The benefit is more on throughput than on response time.
To reduce confusion between the words processor and microprocessor, companies

refer to processors are referred as cores, and such microprocessors are generically
called multicore microprocessors.
Quad core- Microprocessor is a chip that contains four processors or four cores.
1.7 INSTRUCTIONS
Introduction
The repertoire of instructions of a computer
Different computers have different instruction sets
But with many aspects in common
Early computers had very simple instruction sets
Simplified implementation
Many modern computers also have simple instruction sets
The words of a computers language are called instructions, and its vocabulary is called
an instruction set.
Modern instruction set architectures:IA-32, PowerPC, MIPS, SPARC, ARM, and others.
computers are built on two key principles:

1. Instructions are represented as numbers.
2. Programs are stored in memory to be read or written, just like numbers.
These principles lead to the stored-program concept- The idea that instructions and
data of many types can be stored in memory as numbers, leading to the stored program
computer.
16
An instruction set, or instruction set architecture (ISA), is the part of the computer
architecture related to programming, including the native data types, instructions,
registers, addressing modes, memory architecture, interrupt and exception handling,
and external I/O.
An ISA includes a specification of the set of opcodes (machine language), and
the native commands implemented by a particular processor
Classification of instruction sets
A complex instruction set computer (CISC) has many specialized instructions,
which may only be rarely used in practical programs.
A reduced instruction set computer (RISC) simplifies the processor by only

implementing instructions that are frequently used in programs; unusual operations are
implemented as subroutines, where the extra processor execution time is offset by their
rare use.
Basic Instruction Cycle
Elements of Instructions
Operation Code: Specifies the operation to be performed by binary codes. Simply
called opcodes.
Source/Destination operand: The source/destination operand field directly specifies
the source/destination for the instruction.
Source operand address: The operation specified by the instruction may require
one or more source operands.
17
Destination operand address: The operation executed by the CPU may produce
result. Usually, the result is stored in destination operand.
Next instruction Address: The next instruction address tells the CPU from where to
fetch the next instruction after completion of execution of current instruction.
Locations for Source and Destination Operands
Processor registers, Main or virtual memory, I/O Device, Immediate Value-The value of
the operand is given in the instruction itself.
Representation of Instructions
Opcode
Operand Address1
Operand Address2
Instruction type
Data handling and memory operations
Set a register to a fixed constant value.
Move data from a memory location to a register, or vice versa. Used to store the
contents of a register, result of a computation, or to retrieve stored data to perform a
computation on it later.
Read and write data from hardware devices.
Arithmetic and logic operations
Add, subtract, multiply, or divide the values of two registers, placing the result in a
register, possibly setting one or more condition codes in a status register.
Perform bitwise
operations,
e.g.,
taking
the conjunction and disjunction of
corresponding bits in a pair of registers, taking the negation of each bit in a register.
Compare two values in registers (for example, to see if one is less, or if they are
equal).
Control flow operations
Branch to another location in the program and execute instructions there.
Conditionally branch to another location if a certain condition holds.
Indirectly branch to another location, while saving the location of the next instruction
as a point to return to (a call).
Classifying ISAs
The major choices are a stack, an accumulator, or a set of registers.
Operands may be named explicitly or implicitly:
The operands in a stack architecture are implicitly on the top of the stack,
in an accumulator architecture one operand is implicitly the accumulator, and general18
purpose register architectures have only explicit operandseither registers or memory

locations.
The explicit operands may be accessed directly from memory or may need to be
first loaded into temporary storage, depending on the class of instruction and choice of
specific instruction.
Stack
Accumulator
Register
Register
(register-memory)
(load-store)
Push A
Load A
Load R1, A
Load R1,A
Push B
Add B
Add R1, B
Load R2, B
Add
Store C
Store C, R1
Add R3, R1, R2
Pop C
Store C, R3
One can access memory as part of any instruction, called register-memory

architecture, and one can access memory only with load and store instructions, called
load-store or register-register architecture.
A third class, not found in machines shipping today, keeps all operands in
memory and is called a memory-memory architecture.
Examples of instruction set
ADD - Add two numbers together.
COMPARE - Compare numbers.
IN - Input information from a device, e.g. keyboard.
JUMP - Jump to designated RAM address.
JUMP IF - Conditional statement that jumps to a designated RAM address.
LOAD - Load information from RAM to the CPU.
OUT - Output information to device, e.g. monitor.
STORE - Store information to RAM.
MIPS (Million instructions per second)
A measurement of program execution speed based on the number of millions of

instructions.
MIPS is computed as the instruction count divided by the product of the execution time
and 106.
---------- (2.1.1)
19
MIPS is an instruction execution rate, MI PS specifies performance inversely to

execution time; faster computers have a higher MIPS rating.
MIPS Design Principles

1. Simplicity favors regularity
Fixed size instructions
Small number of instruction formats
Opcode always the first 6 bits
2. Smaller is faster
Limited instruction set
Limited number of registers in register file
Limited number of addressing modes
3. Make the common case fast
Arithmetic operands taken only from the register file (load store machine)
Allow instructions to contain immediate operands
4. Good design demands good compromises
only three instruction formats
MIPS Arithmetic Instructions
MIPS Instruction Fields
20
1.8 Operations of the computer hardware
The MIPS assembly language notation instructs a computer to add the two variables b
and c and to put their sum in a.
add a, b, c
This notation is rigid in that each MI PS arithmetic instruction performs only one
operation and must always have exactly three variables.
The following sequence of instructions adds the four variables:

Add a, b, c # The sum of b and c is placed in a.
Add a, b, c # The sum of b,c and d is now in a.
Add a, b, c # The sum of b,c ,d and e is now in a.
Thus, it takes three instructions to sum the four variables.
The words to the right of the sharp symbol (#) on each line above are comments for the
human reader, and the computer ignores them.
Note that unlike other programming languages, each line of this language can contain at
most one instruction.
Another difference from C is that comments always terminate at the end of a line.
The translation from C to MIPS assembly language instructions is performed by the

compiler.
MIPS Operands
Name
Example
Comments
32 registers
$s0$s7, $t0$t9,
Fast locations for data. In MIPS, data must be in
$zero,
registers to perform arithmetic, register $zero
$a0$a3, $v0$v1,
always equals 0, and register $at is reserved by

21
$gp,
the assembler to handle large constants.
$fp,$sp, $ra, $at

230 memory
Memory[0],Memory[4], Accessed only by data transfer instructions. MIPS
words
..,
uses byte addresses, so sequential word
Memory[4294967292]
addresses differ by 4. Memory holds data

structures, arrays, and spilled registers.
MIPS assembly language

Category
Instruction
Add
Arithmetic
Subtract
Meaning
Comments
$s1 = $s2 + $s3
Three register operands
$s1 = $s2 $s3
Three register operands
$s1 = $s2 + 20
Used to add constants
$s1 = Memory[$s2 +
Word from memory to
20]
register
sw
Memory[$s2 + 20] =
Word from register to
$s1,20($s2)
$s1
memory
$s1 = Memory[$s2
Halfword memory to
+20]
register
add
$s1,$s2,$s3
sub
$s1,$s2,$s3
add
addi
immediate
$s1,$s2,20
load word
lw $s1,20($s2)
store word
Data
Example
load half
lh $s1,20($s2)
load half
lhu
$s1 = Memory[$s2
Halfword memory to
unsigned
$s1,20($s2)
+20]
register
sh
Memory[$s2 + 20] =
Halfword register to
$s1,20($s2)
$s1
memory
$s1 = Memory[$s2
Byte from memory to
+20]
register
store half
Transfer
load byte
lb $s1,20($s2)
load byte
lbu
$s1 = Memory[$s2 +
unsigned
$s1,20($s2)
20]
sb
Memory[$s2 + 20] =
Byte from register to
$s1,20($s2)
$s1
memory
$s1 = Memory[$s2 +
Load word as 1st half of
20]
atomic swap
store byte
load linked
word
ll $s1,20($s2)
22
Byte from memory to

register
Store
condition.
sc $s1,20($s2)
word
load upper
immed
and
Or
nor
Logical
lui $s1,20
and
$s1,$s2,$s3
nor
$s1,$s2,$s3
andi
immediate
$s1,$s2,20
immediate
shift left
logical
shift right
$s1=0 or 1
atomic swap
. $s1 = 20 * 216
$s1 = $s2 & $s3
or $s1,$s2,$s3 $s1 = $s2 | $s3
and
or
Memory[$s2+20]=$s1; Store word as 2nd half of
$s1 = ~ ($s2 | $s3)
$s1 = $s2 & 20
Loads constant in upper

16 bits
Three reg. operands; bitby-bit AND
Three reg. operands; bitby-bit OR
Three reg. operands; bitby-bit NOR
Bit-by-bit AND reg with
constant
Bit-by-bit OR reg with
ori $s1,$s2,20
$s1 = $s2 | 20
sll $s1,$s2,10
$s1 = $s2 << 10
srl $s1,$s2,10
$s1 = $s2 >> 10
constant
Shift left by constant
logical Shift right by
constant
Conditional Branch
Instruction
Example
Meaning
Branch on equal
beq $s1,$s2,$s3 If($s1==$s2)
Comments
go
PC+4+100
Branch on not equal
Sent on less than
Sent on less than
bne $s1,$s2,25
slt $s1,$s2,$s3
sltu $s1,$s2,$s3
unsigned
Sent on less than
slt1 $s1,$s2,$s3
immediate
Sent on less than
sltfu
If($s1==$s2)
to Equal test, PC-relative

branch
go
to Not
Equal
test,
PC-
PC+4+100
relative branch
If($s2<$s3)$s1=1;
Compare less than for
Else $s1=0
beq,bne
If($s2<$s3)$s1=1;
Compare less than for
Else $s1=0
unsigned
If($s2<$s3)$s1=1;
Compare
Else $s1=0
constant
If($s2<$s3)$s1=1;
Compare
23
less
than
less
than
immediate
$s1,$s2,$s3
Else $s1=0
constant unsigned
unsigned
Unconditional Branch
Comments
Instruction
Example
Meaning
Jump
j 2500
Go to 10000
Jump to target address
Jump register
jr $ra
Go to $ra
For
switch,
procedure
returns
Jump and link
ja1 2500
$ra=PC+4;go
to For procedure call
10000
1.9 Operands of the Computer Hardware

The operands of arithmetic instructions are restricted; they must be from a limited
number of special locations built directly in hardware called registers.
Registers are the bricks of computer construction: registers are primitives used in
hardware design that are also visible to the programmer when the computer is completed.
The size of a register in the MIPS architecture is 32 bits; groups of 32 bits occur so
frequently that they are given the name word in the MIPS architecture.
One major difference between the variables of a programming language and registers is
the limited number of registers, typically 32 on current computers. MIPS has 32 registers.
Thus, continuing in our top-down, stepwise evolution of the symbolic representation of
the MIPS language, in this section we have added the restriction that the three operands of
MIPS arithmetic instructions must each be chosen from one of the 32-bit registers.
The reason for the limit of 32 registers may be found in the second of our four underlying
design principles of hardware technology:
Design Principle 2: Smaller is faster.
Although we could simply write instructions using numbers for registers, from 0 to 31,
the MIPS convention is to use two-character names following a dollar sign to represent a
register..
24
For now, we will use $s0, $s1, . . . for registers that correspond to variables in C and Java
programs and $t0, $t1, . . . for temporary registers needed to compile the program into MIPS
instructions.
1. Compiling a C Assignment Using Registers
It is the compilers job to associate program variables with registers. Take, for instance,
the assignment statement from our earlier example:
f = (g + h) (i + j);
The variables f, g, h, i, and j are assigned to the registers $s0, $s1, $s2, $s3, and $s4,
respectively.
What is the compiled MIPS code?
The compiled program is very similar to the prior example, except we replace the
variables with the register names mentioned above plus two temporary registers, $t0 and $t1,
which correspond to the temporary variables above:
add $t0,$s1,$s2 # register $t0 contains g + h
add $t1,$s3,$s4 # register $t1 contains i + j
sub $s0,$t0,$t1 # f gets $t0 $t1, which is (g + h)(i + j)
2. Memory Operands
Computer associates variables with registers.
The programs contains lots of variables and more complex data structures like arrays
and structures.
These complex data structures can contain many more data elements than there are
registers in a computer.
However, the processor can keep only a small amount of data in registers, but computer
memory contains billions of data elements. Hence, data structure are kept in memory.
Generally, registers are faster to access than memory and have high throughput than
memory. Accessing register also uses less energy than accessing memory. To achieve highest
performance and conserve energy, compiler must use register efficiently. Therefore, most
frequently used data must placed in register and rest of data into the memory is called spilling
register called stack.
MIPS arithmetic operations occur only on registers in MIPS instructions; thus, MIPS
must include instructions that transfer data between memory and registers. Such instructions
are called Data transfer instructions.
To access a word in memory, the instruction must supply the memory address. Memory
is just a large, single-dimensional array, with the address acting as the index to that array,
starting at 0.
25
The data transfer instruction that copies data from memory to a register is traditionally
called load.
The format of the load instruction is the name of the operation followed by the register to
be loaded, then a constant and register used to access memory.
The sum of the constant portion of the instruction and the contents of the second
register forms the memory address. The actual MIPS name for this instruction is lw, standing
for load word.
3. Compiling an Assignment When an Operand Is in Memory
Eg: Lets assume that A is an array of 100 words and that the compiler has associated the
variables g and h with the registers $s1 and $s2 as before. Let's also assume that the starting
address, or base address, of the array is in $s3. What is the MIPS assembly code for the C
assignment statement?
g = h + A[8];
MIPS Code
In the given statement, there is a single operation. Whereas, one of the operands is in
memory, so must carry this operation in two steps:
Step 1:
Load the temporary register($t0) with memory operand, from the memory
address
A($s3)+8
Step 2:
Perform addition with h($s2) and store result in g($s1)

Lw $t0,8($s3);
#temporary reg $t0 gets A[8](($s3)+8)
The following instruction can operate on the value in $t0 (which equals A[8]) since it is in a
register. The instruction must add h (contained in $s2) to A[8] ($t0) and put the sum in the
register corresponding to g (associated with $s1):
add $s1,$s2,$t0
# g = h + A[8]
Note: The constant in a data transfer instruction is called the offset, and the register added to
form the address is called the base register.
In MIPS, words must start at addresses that are multiples of 4. This requirement is
called an alignment restriction.
Computers divide into those that use the address of the leftmost or big end byte as the word
address versus those that use the rightmost or little end byte.
MIPS is in the Big Endian camp. Byte addressing also affects the array index. To get the
proper byte address in the code above, the offset to be added to the base register $s3 must be
4 8, or 32, so that the load address will select A[8] and not A[8/4].
26
4. Compiling Using Load and Store

Assume variable h is associated with register $s2 and the base address of the array A is in
$s3. What is the MIPS assembly code for the C assignment statement below?
A[12] = h + A[8];
Although there is a single operation in the C statement, now two of the operands are in
memory,so we need even more MIPS instructions.
The first two instructions are the same as the prior example, except this time we use the proper
offset for byte addressing in the load word instruction to select A[8], and the add instruction
places the sum in $t0:
lw $t0,32($s3)
# Temporary reg $t0 gets A[8]
add $t0,$s2,$t0
# Temporary reg $t0 gets h + A[8]
The final instruction stores the sum into A[12], using 48 as the offset and register $s3 as the
base register.
sw $t0,48($s3)
# Stores h + A[8] back into A[12](12*4=48)
5. Constant or Immediate Operands

Constant is a designed value which will not be changed and also not allowed to change
during program execution. More than 50% of MIPS arithmetic and logical instructions have
a constant as an operand i.e., It including constant as a part of the instruction, called
immediate operand.
The hardware principle relating size and speed suggests that memory must be slower
than registers since registers are smaller. This is indeed the case; data accesses are faster if
data is in registers instead of memory.
A MIPS arithmetic instruction can read two registers, operate on them, and write the
result. A MIPS data transfer instruction only reads one operand or writes one operand, without
operating on it.
Thus, MIPS registers take both less time to access and have higher throughput than
memorya rare combinationmaking data in registers both faster to access and simpler to
use. To achieve highest performance, compilers must use registers efficiently.
Using only the instructions we have seen so far, we would have to load a constant from
memory to use one.
For example, to add the constant 4 to register $s3, we could use the code
lw $t0, AddrConstant4($s1) # $t0 = constant 4
add $s3,$s3,$t0 # $s3 = $s3 + $t0 ($t0 == 4)
assuming that AddrConstant4 is the memory address of the constant 4.
27
An alternative that avoids the load instruction is to offer versions of the arithmetic
instructions in which one operand is a constant. This quick add instruction with one constant
operand is called add immediate or addi. To add 4 to register $s3, we just write
addi $s3,$s3,4 # $s3 = $s3 + 4
Immediate instructions illustrate the third hardware design principle, first mentioned in the
Fallacies and Pitfalls of Chapter 1:
Design Principle 3: Make the common case fast.
Constant operands occur frequently, and by including constants inside arithmetic
instructions, they are much faster than if constants were loaded from memory.
Importance of registers
What is the rate of increase in the number of registers in a chip over time?
1. Very fast: They increase as fast as Moores law, which predicts doubling the number
of transistors on a chip every 18 months.
2. Very slow: Since programs are usually distributed in the language of the computer,
there is inertia in instruction set architecture, and so the number of registers increases
only as fast as new instruction sets become viable.
1.10
Representing Instructions in the Computer
Computers are built on two key principles:

1. Instructions are represented as numbers.
2. Programs are stored in memory to be read or written, just like numbers.
Instructions are also kept in the computer as a series of high and low electronic signals
and may be represented as numbers.
In fact, each piece of an instruction can be considered as an individual number, and
placing these numbers side by side forms the instruction.
Since registers are part of almost all instructions, there must be a convention to map
register names into numbers. In MIPS assembly language, registers $s0 to $s7 map onto
registers 16 to 23, and registers $t0 to $t7 map onto registers 8 to 15.
Hence, $s0 means register 16, $s1 means register 17, $s2 means register 18, . . . , $t0
means register 8, $t1 means register 9, and so on.
1. Translating a MIPS Assembly Instruction into a Machine Instruction
Real MIPS language version of the instruction represented symbolically as
add $t0,$s1,$s2
first as a combination of decimal numbers and then of binary numbers.
The decimal representation is
28
17
18
32
Each of these segments of an instruction is called a field.
The first and last fields (containing 0 and 32 in this case) in combination tell the MIPS
computer that this instruction performs addition.
The second field gives the number of the register that is the first source operand of the
addition operation (17 = $s1).
The third field gives the other source operand for the addition (18 = $s2).
The fourth field contains the number of the register that is to receive the sum (8 = $t0).
The fifth field is unused in this instruction, so it is set to 0.
Thus, this instruction adds register $s1 to register $s2 and places the sum in register $t0.
This instruction can also be represented as fields of binary numbers as opposed to decimal:
000000
6 bits
10001
5 bits
10010
01000
5 bits
00000
5 bits
5 bits
100000
6 bits
MIPS Fields
MIPS fields are given names to make them easier to discuss:
op: Basic operation of the instruction, traditionally called the opcode.
rs: The first register source operand.
rt: The second register source operand.
rd: The register destination operand. It gets the result of the operation.
shamt: Shift amount. It will not be used until then, and hence the field contains zero.
funct: Function. This field selects the specific variant of the operation in the op field
and is sometimes called the function code.
A problem occurs when an instruction needs longer fields than those shown above. For
example, the load word instruction must specify two registers and a constant. If the address
were to use one of the 5-bit fields in the format above, the constant within the load word
instruction would be limited to only 25 or 32. This constant is used to select elements from
arrays or data structures, and it often needs to be much larger than 32. This 5-bit field is too
small to be useful.
Hence, we have a conflict between the desire to keep all instructions the same length and the
desire to have a single instruction format. This leads us to the final hardware design principle:
Design Principle 4: Good design demands good compromises.
29
The compromise chosen by the MIPS designers is to keep all instructions the same length,
thereby requiring different kinds of instruction formats for different kinds of instructions.
2. Translating MIPS Assembly Language into Machine Language
If $t1 has the base of the array A and $s2 corresponds to h, the assignment statement
A[300] = h + A[300];
is compiled into
lw $t0,1200($t1)
# Temporary reg $t0 gets A[300]
add $t0,$s2,$t0
# Temporary reg $t0 gets h + A[300]
sw $t0,1200($t1)
# Stores h + A[300] back into A[300]
What is the MIPS machine language code for these three instructions?
For convenience, represent the machine language instructions using decimal numbers.
op
rs
rt
35
18
43
rd
shamt
function
1200
8
32
1200
1. The lw instruction is identified by 35 in the first field (op).

2. The base register 9 ($t1) is specified in the second field (rs)
3. The destination register 8 ($t0) is specified in the third field (rt).
4. The offset to select A[300] (1200 = 300 4) is found in the final field (address).
The add instruction that follows is specified with 0 in the first field (op) and 32 in the last field
(funct). The three register operands (18, 8, and 8) are found in the second, third, and fourth
fields and correspond to $s2, $t0, and $t0.
The sw instruction is identified with 43 in the first field. The rest of this final instruction is
identical to the lw instruction.
The binary equivalent to the decimal form is the following (1200 in base 10 is 0000 0100 1011
0000 base 2):
Why doesnt MIPS have a subtract immediate instruction?

1. Negative constants appear much less frequently in C and Java, so they are not the
common case and do not merit special support.
30
2. Since the immediate field holds both negative and positive constants, add immediate
with a negative number is equivalent to subtract immediate with a positive number, so subtract
immediate is superfluous.
MIPS Instruction Coding Format
MIPS instructions are classified into 4 groups according to their coding formats:
1. R-type (for register) or R-format
2. I-type (for immediate) or I-format
3. J-type(for Jump) or J-format
1. R-type (for register) or R-format
MIPS Instruction encoding
MIPS = RISC hence
Few (3+) instruction formats
R in RISC also stands for Regular
All instructions of the same length (32-bits = 4 bytes)
Formats are consistent with each other
Opcode always at the same place (6 most significant bits)
rd and rs always at the same place
immed always at the same place etc.
Opcode(6)
rs(5)
rt(5)
rd(5)
sa(5)
2. I-type (for immediate) or I-format

An instruction with an immediate constant has
Opcode(6)
rs(5)
rt(5)
Immediate or address(16)
Encoding of the 32 bits:

Opcode is 6 bits
Since we have 32 registers, each register name is 5 bits
This leaves 16 bits for the immediate constant
I-type format example
Addi $a0,$12,33
# $a0 is also $4 = $12 +33
# Addi has opcode 08
Opcode(6)
08
rs(5)
12
rt(5)
4
Immediate or address(16)
33
31
Function(6)
1.11 Logical Operations

Logical Operations are useful for extracting and inserting groups of bits in a word and it
needs to operate on fields of bits within a word or even on individual bits.
Examining characters which are stored as bytes within words
Packing and unpacking of bits into words

Logical
C operators
Java operators
MIPS instructions
operations
Shift left
<<
<<
sll
Shift right
>>
>>>
srl
Bit-by-bit AND
&
&
and,andi
Bit-by-bit OR
or, ori
Bit-by-bit NOT
nor
1. Shift left
The first class operation is called shifts. They move all the bits in a word to the left or
right, filling the emptied bits with 0s. The actual name of the MIPS shift instructions are called
shift left logical (sll)
For example, if register $s0 contained
=
and the instruction to shift left by 4 was executed, the new value would look like this:
2. Shift right
The dual of a shift left is a shift right. The actual name of the MIPS shift instructions are
called shift right logical (srl).
The following instruction performs the operation above, assuming that the result should
go in register $t2:
sll $t2,$s0,4
# reg $t2 = reg $s0 << 4 bits
We delayed explaining the shamt field in the R-format. It stands for shift amount and is used in
shift instructions. Hence, the machine language version of the instruction above is
op
rs
rt
rd
32
shamt
funct
16
10
The encoding of sll is 0 in both the op and funct fields, rd contains $t2, rt contains $s0,
and shamt contains 4. The rs field is unused, and thus is set to 0.
Shift left logical provides a bonus benefit. Shifting left by i bits gives the same result as
multiplying by 2i.
3. Bit-by-bit AND
Another useful operation that isolates fields is AND. AND is a bit-by-bit operation that
leaves a 1 in the result only if both bits of the operands are 1.
For example, if register $t2 still contains
0000 0000 0000 0000 0000 1101 0000 0000two
and register $t1 contains
0000 0000 0000 0000 0011 1100 0000 0000two
then, after executing the MIPS instruction
and $t0,$t1,$t2 # reg $t0 = reg $t1 & reg $t2 the value of register $t0 would be
0000 0000 0000 0000 0000 1100 0000 0000two
AND can apply a bit pattern to a set of bits to force 0s where there is a 0 in the bit
pattern. Such a bit pattern in conjunction with AND is traditionally called a mask, since the
mask conceals some bits.
4. Bit-by-bit OR
To place a value into one of these seas of 0s, there is the dual to AND, called OR. It is a bitby-bit operation that places a 1 in the result if either operand bit is a 1.
To elaborate, if the registers $t1 and $t2 are unchanged from the preceding example, the
result of the MIPS instruction
or $t0,$t1,$t2
# reg $t0 = reg $t1 | reg $t2
is this value in register $t0:

0000 0000 0000 0000 0011 1101 0000 0000two
5. Bit-by-bit NOT
The final logical operation is a contrarian.
NOT takes one operand and places a 1 in the result if one operand bit is a 0, and vice
versa.
In keeping with the two operand format, the designers of MIPS decided to include the
instruction NOR (NOT OR) instead of NOT.
33
If one operand is zero, then it is equivalent to NOT.
For example, A NOR 0 = NOT (A OR 0) = NOT (A).

If the register $t1 is unchanged from the preceding example and register $t3 has the value 0,
the result of the MIPS instruction
nor $t0,$t1,$t3 # reg $t0 = ~ (reg $t1 | reg $t3)
is this value in register $t0:
11111111 1111 1111 1100 0011 1111 1111two
1.12 Control Operation

Decision making instructions alter the control flow, i.e, change the next instruction to
be executed, like if statement with a goto statement.
MIPS assembly language includes two decision-making instructions, similar to an if-else
and go-to called conditional branch instruction.
First instruction: beq stands for branch if equal
beq register1, register2, L1;
Branch if equal statement state that if value of register1 is equal to the value of
register2, then processor can change its sequence to L1, otherwise it execute the next
instruction in sequence.
This instruction means go to the statement labeled L1 if the value in register1 equals the
value in register2.
Second instruction: bne stands for branch if not equal
bne register1, register2, L1
It means go to the statement labeled L1 if the value in register1 does not equal the value
in register2.
These two instructions are traditionally called conditional branches.
1. Compiling if-then-else into Conditional Branches
In the following code segment, f, g, h, i, and j are variables. If the five variables f through j
correspond to the five registers $s0 through $s4.
What is the compiled MIPS code for this C if statement?
if (i == j) f = g + h; else f = g h;
The first expression compares for equality, so it would seem that we would want beq.
In general, the code will be more efficient if we test for the opposite condition to branch
over the code that performs the subsequent then part of the if :
bne $s3,$s4,Else
# go to Else if i # j
34
The next assignment statement performs a single operation, and if all the
operands are allocated to registers, it is just one instruction:
add $s0,$s1,$s2
# f = g + h (skipped if i #j)
Unconditional branch
This instruction says that the processor always follows the branch. To distinguish
between conditional and unconditional branches, the MIPS name for this type of instruction is
jump, abbreviated as j.
j Exit
# go to Exit
The assignment statement in the else portion of the if statement can again be compiled
into a single instruction.
We just need to append the label Else to this instruction. We also show the label Exit
that is after this instruction, showing the end of the if-then-else compiled code:
Else:sub $s0,$s1,$s2
# f = g h (skipped if i #j)
Exit:
Loops
2. Compiling a while Loop in C
Traditional loop in C:
while (save[i] == k)
i += 1;
Assume:
i and k correspond to registers $s3 and $s5 and the base of the array save is in
$s6.
What is the MIPS assembly code corresponding to this C segment?
1. To load save[i] into a temporary register.
2. Before loading save[i] into a temporary register, we need to have its address.
3. Before adding i to the base of array save to form the address, we must multiply the
index i by 4 due to the byte addressing problem.
4. Use shift left logical since shifting left by 2 bits multiplies by 4.
5. Add the label Loop to it so that we can branch back to that instruction at the end of
the loop:
Loop: sll $t1,$s3,2
# Temp reg $t1 = 4 * i
To get the address of save[i], we need to add $t1 and the base of save in $s6:
add $t1,$t1,$s6
# $t1 = address of save[i]
Now we can use that address to load save[i] into a temporary register:
lw $t0,0($t1)
# Temp reg $t0 = save[i]

35
The next instruction performs the loop test, exiting if save[i] # k:

bne $t0,$s5, Exit
# go to Exit if save[i]
The next instruction adds 1 to i:

add $s3,$s3,1
#i=i+1
The end of the loop branches back to the while test at the top of the loop.
Add the Exit label
j Loop
# go to Loop
Exit:
The test for equality or inequality is probably the most popular test, but sometimes it is
useful to see if a variable is less than another variable.
For example, a for loop may want to test to see if the index variable is less than 0. Such
comparisons are accomplished in MIPS assembly language with an instruction that compares
two registers and sets a third register to 1 if the first is less than the second; otherwise, it is set
to 0.
The MIPS instruction is called set on less than, or slt.
For example,
slt $t0, $s3, $s4
means that register $t0 is set to 1 if the value in register $s3 is less than the value in register
$s4; otherwise, register $t0 is set to 0.
Constant operands are popular in comparisons. Since register $zero always has 0, we
can already compare to 0. To compare to other values, there is an immediate version of the set
on less than instruction. To test if register $s2 is less than the constant 10, we can just write
slti $t0,$s2,10
# $t0 = 1 if $s2 < 10
MIPS compilers use the slt, slti, beq, bne, and the fixed value of 0 to create all relative
conditions: equal, not equal, less than, less than or equal, greater than, greater than or equal.
3. Case/Switch Statement
A case or switch statement that allows the programmer to select one of many
alternatives depending on a single value.
The simplest way to implement switch is via a sequence of conditional tests, turning the
switch statement into a chain of if-then-else statements.
Sometimes the alternatives may be more efficiently encoded as a table of addresses of

alternative instruction sequences, called a jump address table, and the program needs
only to index into the table and then jump to the appropriate sequence.
36
The jump table is then just an array of words containing addresses that correspond to
labels in the code.
To support such situations, computers like MIPS include a jump register instruction (jr),
meaning an unconditional jump to the address specified in a register.
The program loads the appropriate entry from the jump table into a register, and then it
jumps to the proper address using a jump register.
4. Supporting Procedures in Computer Hardware

A procedure or function is one tool C or Java programmers use to structure programs,
both to make them easier to understand and to allow code to be reused.
Procedures allow the programmer to concentrate on just one portion of the task at a
time, with parameters acting as a barrier between the procedure and the rest of the program
and data, allowing it to be passed values and return results.
Similarly, in the execution of a procedure, the program must follow these six steps:
1. Place parameters in a place where the procedure can access them.
2. Transfer control to the procedure.
3. Acquire the storage resources needed for the procedure.
4. Perform the desired task.
5. Place the result value in a place where the calling program can access it.
6. Return control to the point of origin, since a procedure can be called from several points in a
program.
MIPS software follows the following convention in allocating its 32 registers for procedure
calling:
$a0$a3: four argument registers in which to pass parameters
$v0$v1: two value registers in which to return values
$ra: one return address register to return to the point of origin
The jump-and-link instruction (jal) is simply written
jal ProcedureAddress
The link portion of the name means that an address or link is formed that points to the
calling site to allow the procedure to return to the proper address.
This link, stored in register $ra, is called the return address.
The return address is needed because the same procedure could be called from several
parts of the program.
The jal instruction saves PC + 4 in register $ra to link to the following instruction to set
up the procedure return.
37
To support such situations, computers like MIPS use a jump register instruction (jr),
meaning an unconditional jump to the address specified in a register:
jr $ra
The jump register instruction jumps to the address stored in register $rawhich is just
what we want. Thus, the calling program, or caller, puts the parameter values in $a0$a3 and
uses jal X to jump to procedure X (sometimes named the callee). The callee then performs the
calculations, places the results in $v0$v1, and returns control to the caller using jr $ra.
1.13
Addressing Modes
The addressing mode specifies a rule for interpreting or modifying the address field of
the instruction before the operand is actually referenced.

Memory may be viewed as a single-dimensional array of individually addressable bytes. 32bit words are aligned to 4 byte boundaries.
232 bytes, with addresses from 0 to 232 1
230 words with addresses 0, 4, 8, , 232 - 4
Use of addressing modes in computers
Computers use addressing mode techniques for the purpose of accommodating one or both of
the following:
1. To give programming versatility to the user by providing such facilities as pointers to
memory, counters for loop control, indexing of data and program relocation.
2. To reduce the number of bits in the addressing field of the instruction.
Memory Addressing
Independent of whether the architecture is register-register or allows any operand to be
a memory reference, it must define how memory addresses are interpreted and how they are
specified.
Interpreting Memory Addresses
How is a memory address interpreted? That is, what object is accessed as a function of
the address and the length? All the instruction sets are byte addressed and provide access for
bytes (8 bits), half words (16 bits), and words (32 bits). Most of the machines also provide
access for double words (64 bits).
S.No
Little Endian
Big Endian
38
1.
The address of a datum is the address The address of a datum is the

of the least-significant byte.
address of the most-significant byte
Little Endian byte order puts the byte Big Endian byte order puts the byte
2.
whose address is x...x00 at the least whose address is x...x00 at the

significant position in the word (the most-significant position in the word
little end)
(the big end).
Ex: 0000 0001 0010 0011 0100 0101 0110 0111

Little Endian
x..x1101
x..x1100
x..x0111
00000001
x..x0110
00100011
x..x0101
01000101
x..x0100
01100111
x..x0011
x..x0010
Big Indian
x..x1101
x..x1100
x..x0111
01100111
x..x0110
01000101
x..x0101
00100011
x..x0100
00000001
x..x0011
x..x0010
39
When operating within one machine, the byte order is often unnoticeableonly
programs that access the same locations as both words and bytes can notice the difference.
MIPS Addressing Modes:
Multiple forms of addressing are generically called addressing modes. The MIPS addressing
modes are the following:
1. Register addressing
Operands are in a register.
Example: add $3,$4,$5
Takes n bits to address 2n registers
2. Base or displacement addressing, where the operand is at the memory location whose
address is the sum of a register and a constant in the instruction.
The address of the operand is the sum of the immediate and the value in a register (rs).
16-bit immediate is a twos complement number
Ex: lw $15,16($12)
3. Immediate addressing, where the operand is a constant within the instruction itself.
16 bit twos-complement number:
-215 1 = -32,769 < value < +215 = +32,768
4. PC-relative addressing, where the address is the sum of the PC and a constant in the
instruction. The value in the immediate field is interpreted as an offset of the next instruction
(PC+4 of current instruction)
Example: beq $0,$3,Label
5. Pseudo direct addressing, where the jump address is the 26 bits of the instruction
concatenated with the upper bits of the PC Note that a single operation can use more
than one addressing mode.
Add, for example, uses both immediate (addi) and register (add) addressing.
26 bits of the address is embedded as the immediate, and is used as the instruction offset
within the current 256MB (64MWord) region defined by the MS 4 bits of the PC.
Example: j Label
1. Immediate addressing
2. Register addressing
40
3. Base addressing
4. PC-relative addressing
5. Pseudo direct addressing
The words of a computer's language are called instructions, and its vocabulary is called an
instruction set. that operates on them.
It is easy to see by formal-logical methods that there exist certain [instruction sets] that are in
abstract adequate to control and cause the execution of any sequence of operations. . . . The
really decisive considerations from the present point of view, in selecting an [instruction set],
are more of a practical nature: simplicity of the equipment demanded by the [instruction set],
and the clarity of its application to the actually important problems together with the speed of its
handling of those problems.
The chosen instruction set comes from MIPS, which is typical of instruction sets designed
since the 1980s. Almost 100 million of these popular microprocessors were manufactured in
2002, and they are found in products from ATI Technologies, Broadcom, Cisco, NEC,
Nintendo, Silicon Graphics, Sony, Texas Instruments, and Toshiba, among others.
41

UNIT 1 Ca

Caricato da

Informazioni sul documento

Titolo originale

Copyright

Formati disponibili

Condividi questo documento

Condividi o incorpora il documento

Opzioni di condivisione

Hai trovato utile questo documento?

Questo contenuto è inappropriato?

Copyright:

Formati disponibili

UNIT 1 Ca

Caricato da

Copyright:

Formati disponibili

P. R.

Department of CSE & IT

DEPARTMENT OF CSE AND IT

LECTURE NOTES AND QUESTION BANK FOR

CS6303 COMPUTER ARCHITECTURE

Department of CSE & IT

UNIT I OVERVIEW & INSTRUCTIONS 9

Department of CSE & IT

UNIT I - OVERVIEW & INSTRUCTIONS

Eight Great Ideas

Design for Moores Law

It states that integrated circuit resources double every 1824 months.

1.1.2 Use Abstraction to Simplify Design

A major productivity technique for hardware and software is to use abstractions.

To represent the design at different levels of representation; lower-levels and higher

1.1.3 Make the Common Case Fast

Performance via Parallelism

Performance via Pipelining

Department of CSE & IT

A particular pattern of parallelism is pipelining.

Figure 1.1 Memory hierarchy

Dependability via Redundancy

Computers not only need to be fast; they need to be dependable.

1.2 COMPONENTS OF A COMPUTER SYSTEM (Functional Units)

A modern computer consisting of hardware and software.

Department of CSE & IT

Datapath unit - It performs the arithmetic operations.

Figure 1.2: Components of computer system

Department of CSE & IT

It stores the program and data.

Department of CSE & IT

The image is composed of a matrix of picture elements or pixels, which can be

1.3 TECHNOLOGIES FOR BUILDING PROCESSORS AND MEMORY

The technology shapes what computers will be able to do and how

Integrated circuit (IC)

Very large scale IC (VLSI)

Ultra large scale IC(ULSI)

A transistor is simply an on/off switch controlled by electricity.

Department of CSE & IT

Very Large-Scale Integrated (VLSI) circuit-A device containing hundreds of

Integrated Circuit Cost

Department of CSE & IT

Ultra Large-Scale Integrated

To maximize performance, we want to minimize response time or execution time for

Computer performance evaluation is primarily based on throughput and response time.

The manager of a large data processing center may be interested in increasing

To maximize performance, we want to minimize response time or execution time for

The relation between performance and execution time for a computer X:

If X is n times as fast as Y, then the execution time on Y is n times as long as it is on X:

Department of CSE & IT

Table 1: The basic components of performance and its units of measure.

CPU execution time for a program

Seconds for the program

Instructions Executed for the Program

Clock cycles per instruction (CPI)

Average number of clock cycles per instruction

Clock cycle time

Seconds per clock cycle

The Classic CPU Performance Equation

Measuring Execution Time

Department of CSE & IT

_ How do we compare one computer/processor against another?

Department of CSE & IT