| title | Embedded Architectures |
|---|---|
| textbook | #embedded-architectures |
- like autonomy, no exact definition
- like autonomy, no exact definition
- systems created for specific functionality
- like autonomy, no exact definition
- systems created for specific functionality
- fixed/operational for years, decades even!
performance vs
performance vs
- power/battery life
performance vs
- power/battery life
- less memory
performance vs
- power/battery life
- less memory
- fewer peripherals
performance vs
- power/battery life
- less memory
- fewer peripherals
- limited applications
performance vs
- power/battery life
- less memory
- fewer peripherals
- limited applications
- smaller operating systems
- etc.
main reason?
guarantee the system works,
guarantee the system works,
- correctly
- safely
guarantee the system works,
- correctly
- safely
essentially → must be easy to certify the system
"worst-case execution time"
"worst-case execution time"
longest execution time for a program
"worst-case execution time"
longest execution time for a program
on a specific hardware platform
"worst-case execution time"
longest execution time for a program
on a specific hardware platform
must consider → all possible inputs
- "schedulability"
- resource requirements
- performance limits
of embedded and real-time programs
lots of approaches to computing the WCET, e.g.,
- dynamic/empirical → run program lots of times on the platform
- static → compile time analysis to compute the worst-case paths
- hybrid → a combination of the two
- probabilistic → combine dynamic analysis+statistical methods
- ML-based methods → applying machine-learning to the problem
at a high-level, the execution time distributions of applications:
Note: we see that the various methods either underestimate or overestimate the wcet. the former affects the safety, the latter wastes resources
structural problems (in hardware/software) make it more difficult!
void main()
{
int max = 10 ;
int sum = 0;
for( int i = 0 ; i < max ; ++i)
sum += i ;
}void main()
{
int max = 10 ;
int sum = 0;
for( int i = 0 ; i < max ; ++i)
sum += i ;
}how do you compute the WCET for this code?
first, define the processor it runs on → say, "P"
first, define the processor it runs on → say, "P"
- how long each instruction takes to execute on P
first, define the processor it runs on → say, "P"
- how long each instruction takes to execute on P
- how many loop iterations?
first, define the processor it runs on → say, "P"
- how long each instruction takes to execute on P
- how many loop iterations?
- what are the startup/cleanup times for the program on P?
void main() // startup cost = 100 cycles
{
int max = 15 ; // 10 cycles
int sum = 0; // 10 cycles
for( int i = 0 ; i < max ; ++i) // 5 cycles, once
sum += i ; // 20 cycles each iteration
} // cleanup cost = 120 cycles1 void main() // startup cost = 100 cycles
2 {
3 int max = 15 ; // 10 cycles
4 int sum = 0; // 10 cycles
5 for( int i = 0 ; i < max ; ++i) // 5 cycles, once
6 sum += i ; // 20 cycles each iteration
7 } // cleanup cost = 120 cycles1 void main() // startup cost = 100 cycles
2 {
3 int max = 15 ; // 10 cycles
4 int sum = 0; // 10 cycles
5 for( int i = 0 ; i < max ; ++i) // 5 cycles, once
6 sum += i ; // 20 cycles each iteration
7 } // cleanup cost = 120 cycleswhich is fine for this simple example
Now consider this slight change to the above code:
void main( int argc, char* argv[] )
{
int max = atoi( argv[1] ) ; // convert the command line arg to max
int sum = 0;
for( int i = 0 ; i < max ; ++i) // how many iterations?
sum += i ;
}Now consider this slight change to the above code:
void main( int argc, char* argv[] )
{
int max = atoi( argv[1] ) ; // convert the command line arg to max
int sum = 0;
for( int i = 0 ; i < max ; ++i) // how many iterations?
sum += i ;
}previous equation fails!
Now consider this slight change to the above code:
void main( int argc, char* argv[] )
{
int max = atoi( argv[1] ) ; // convert the command line arg to max
int sum = 0;
for( int i = 0 ; i < max ; ++i) // how many iterations?
sum += i ;
}previous equation fails!
no longer know the value of max → arbitrary wcet!
let's make another change to the code:
#define VERY_LARGE_ARRAY+SIZE 1>>18
void main()
{
int first_array[VERY_LARGE_ARRAY_SIZE] ;
int second_array[VERY_LARGE_ARRAY_SIZE] ;
int sum_first = 0;
int sum_second = 0;
for( int i = 0 ; i < VERY_LARGE_ARRAY_SIZE * 2 ; ++i)
{
if( i%2 )
first_sum += first_array[i/2] ;
else
second_sum += second_array[(int)((i/2)+1)] ;
}
}let's make another change to the code:
#define VERY_LARGE_ARRAY+SIZE 1>>18
void main()
{
int first_array[VERY_LARGE_ARRAY_SIZE] ;
int second_array[VERY_LARGE_ARRAY_SIZE] ;
int sum_first = 0;
int sum_second = 0;
for( int i = 0 ; i < VERY_LARGE_ARRAY_SIZE * 2 ; ++i)
{
if( i%2 )
first_sum += first_array[i/2] ;
else
second_sum += second_array[(int)((i/2)+1)] ;
}
}can we compute wcet easily (using previous equation)?
note: maximum size of loop is known → VERY_LARGE_ARRAY_SIZE * 2
Note: pause and ask students what, if anything can go wrong?
| iter | operation | cache state | reason |
|---|---|---|---|
| 1 | first_array loaded |
miss | evicts whatever was previously in cache |
| iter | operation | cache state | reason |
|---|---|---|---|
| 1 | first_array loaded |
miss | evicts whatever was previously in cache |
| 2 | second_array loaded |
miss | evicts first_array due to lack of space |
| iter | operation | cache state | reason |
|---|---|---|---|
| 1 | first_array loaded |
miss | evicts whatever was previously in cache |
| 2 | second_array loaded |
miss | evicts first_array due to lack of space |
| 3 | first_array loaded again |
miss | evicts second_array due to lack of space |
| ... | |||
program will constantly suffer cache misses!
program will constantly suffer cache misses!
#define VERY_LARGE_ARRAY+SIZE 1>>18
void main()
{
int first_array[VERY_LARGE_ARRAY_SIZE] ;
int second_array[VERY_LARGE_ARRAY_SIZE] ;
int sum_first = 0;
int sum_second = 0;
for( int i = 0 ; i < VERY_LARGE_ARRAY_SIZE * 2 ; ++i)
{
if( i%2 )
first_sum += first_array[i/2] ;
// 20 cycles + cache miss time (100s of cycles)
else
second_sum += second_array[(int)((i/2)+1)] ;
// 20 cycles + cache miss time (100s of cycles)
}
}program will constantly suffer cache misses!
now we need to model → cache behavior for program/data
other hardware features, e.g.,
- processor pipelining
- prefetching
- branch prediction
- multithreading
- multicore systems
- memory buses
- networks-on-chip
- and too many others to recount here...
processor feature → improves performance → bad for wcet analysis!
processor feature → improves performance → bad for wcet analysis!
come in myriad shapes and sizes:
- Microcontrollers
- Digital Signal Processors
- Microprocessors
- System-on-a-Chip
- Embedded accelerators
- ASICs and FPGAs
so, we focus on:
from Wikipedia,
"A microcontroller (MC, UC, or μC) or microcontroller unit (MCU) is a small computer
on a single integrated circuit."
- most common type of "processors" in embedded systems
- more than 55% of the world's processors!
- most common type of "processors" in embedded systems
- more than 55% of the world's processors!
- small, yet critical, systems
- most common type of "processors" in embedded systems
- more than 55% of the world's processors!
- small, yet critical, systems
- car engine control
- implantable medical devices
- thermal monitoring
- millions (billions?) of other applications!
| component | details |
|---|---|
| one CPU core | typically simple 4 or 8 bit chips |
Note: Processor bit width refers to the amount of data (in bits) a CPU can process in a single operation, determining its data handling capacity and memory addressing ability, with 32-bit and 64-bit being common standards, where 64-bit CPUs handle larger chunks of data (8 bytes) and access vastly more RAM than 32-bit ones (4 bytes)
| component | details |
|---|---|
| one CPU core | typically simple 4 or 8 bit chips |
| small pipelined architectures | sometimes 2 or 4 stage pipelines |
| component | details |
|---|---|
| one CPU core | typically simple 4 or 8 bit chips |
| small pipelined architectures | sometimes 2 or 4 stage pipelines |
| some limited memory | typically a few hundred kilobytes, perhaps in the form of EEPROMs or FLASH |
| component | details |
|---|---|
| one CPU core | typically simple 4 or 8 bit chips |
| small pipelined architectures | sometimes 2 or 4 stage pipelines |
| some limited memory | typically a few hundred kilobytes, perhaps in the form of EEPROMs or FLASH |
| programmable I/O | to interact with the real world |
| component | details |
|---|---|
| low operating frequencies | e.g., 4 KHz; simpler/older processors, yet more predictable |
Note: processor "frequency is the rate at which a CPU executes basic operations, measured in gigahertz (GHz), indicating billions of cycles per second, essentially how many "ticks" the processor makes to process instructions;
| component | details |
|---|---|
| low operating frequencies | e.g., 4 KHz; simpler/older processors, yet more predictable |
| low power consumption | in the milliwatts or microwatts ranges; might even be nanowatts when the system is sleeping |
| component | details |
|---|---|
| low operating frequencies | e.g., 4 KHz; simpler/older processors, yet more predictable |
| low power consumption | in the milliwatts or microwatts ranges; might even be nanowatts when the system is sleeping |
| interrupts (some programmable) | often real-time (fixed/low latency) |
Note: this is the crux of how we can implement "real-time" behavior. Low latency/deterministic interrupts
| component | details |
|---|---|
| low operating frequencies | e.g., 4 KHz; simpler/older processors, yet more predictable |
| low power consumption | in the milliwatts or microwatts ranges; might even be nanowatts when the system is sleeping |
| interrupts (some programmable) | often real-time (fixed/low latency) |
| general-purpose I/O (GPIO) pins | for I/O |
| component | details |
|---|---|
| low operating frequencies | e.g., 4 KHz; simpler/older processors, yet more predictable |
| low power consumption | in the milliwatts or microwatts ranges; might even be nanowatts when the system is sleeping |
| interrupts (some programmable) | often real-time (fixed/low latency) |
| several general-purpose I/O (GPIO) pins | for I/O |
| timers | e.g., a programmable interval timer (PIT) |
| component | details |
|---|---|
| analog to digital (ADC) convertors | to convert incoming (real-world, sensor) data to a digital form that the uC can operate on |
Note:
- because the real world is "analog"/continuous
| component | details |
|---|---|
| analog to digital (ADC) convertors | to convert incoming (real-world, sensor) data to a digital form that the uC can operate on |
| digital-to-analog (DAC) convertor | to do the opposite, convert from digital to analog signals to send outputs in that form |
Note:
- because the real world is "analog"/continuous
| component | details |
|---|---|
| analog to digital (ADC) convertors | to convert incoming (real-world, sensor) data to a digital form that the uC can operate on |
| digital-to-analog (DAC) convertor | to do the opposite, convert from digital to analog signals to send outputs in that form |
| universal async tx/rx (UART) | to receive/send data over a serial line |
| component | details |
|---|---|
| pulse width modulation (PWM) | so that the CPU can control motors (significant for us in autonomous/automotive systems), power systems, resistive loads, etc. |
| component | details |
|---|---|
| pulse width modulation (PWM) | so that the CPU can control motors (significant for us in autonomous/automotive systems), power systems, resistive loads, etc. |
| JTAG interface | debugging interface |
- are small → must fit in (limited) memory
- are small → must fit in (limited) memory
- often directly programmed in assembly!
- are small → must fit in (limited) memory
- often directly programmed in assembly!
- assembly code might need hand tuning
Note: hand-tuning is for both, performance as well as fitting into the limited memory
- are small → must fit in (limited) memory
- often directly programmed in assembly!
- assembly code might need hand tuning
- C is another popular language
- are small → must fit in (limited) memory
- often directly programmed in assembly!
- assembly code might need hand tuning
- C is another popular language
- no operating systems (or very rare)!
Note:
- "**bare metal execution"
- are small → must fit in (limited) memory
- often directly programmed in assembly!
- assembly code might need hand tuning
- C is another popular language
- no operating systems (or very rare)!
- sometimes use special-purpose programming languages/instructions
![]() Atmel ATmega |
![]() Microchip Technology |
![]() Motorola (Freescale) |
![]() NXP |
- specialized microcontrollers optimized for digital signal processing
- specialized microcontrollers optimized for digital signal processing
- specialized microcontrollers optimized for digital signal processing
- widely used in:
- audio processing
- radar and sonar
- speech recognition systems
- image processing
- satellites
- telecommunications
- mobile phones
- televisions
- isolate, measure, compress, filter analog signals in the real world
- often have stringent real-time constraints
- texas instruments dsp chip, tms320 series
- one of the most famous examples of dsp systems
- repetitive mathematical operations
- over a large number of samples
- in real-time
- repetitive mathematical operations
- over a large number of samples
- in real-time
Note: Imagine you are reading a sensor that measures temperature every second. The signal is "noisy" (it jumps around erratically). To smooth it, you don't just look at the current number; you look at the average of the current number and the previous two numbers.
This requires two repetitive steps that are the heart of DSP:
Multiplication: Weighting the data points.
Accumulation: Adding them together.
signal = [10, 22, 10, 12]
weights = [0.5, 0.25, 0.25]
output = []
# This loop is the "repetitive" nature of DSP
for i in range(2, len(signal)):
# The MAC (Multiply-Accumulate) Operation
val = (signal[i] * weights[0]) + \
(signal[i-1] * weights[1]) + \
(signal[i-2] * weights[2])
output.append(val)
print(output) # Result: [13.5, 14.0]- analog to digital conversion
- manipulation (the core algorithm)
- digital to analog conversion
- entire process → completed with low latency, fixed deadline
- low power requirements → battery-constrained devices
- examples of software implementations:
- liquid sdr
- matlab's dsp system toolbox
Note: hardware is faster, more predictable, low latency/jitter
- custom instructions
- e.g.,
saturate,ed,multiply-and-accumulate
- e.g.,
Note:
- caps the minimum or maximum value that can be held in a fixed-point representation (e.g., moving audio knob past 100 → shouldn't "wrap around" to 0),
- Euclidean distance, a+b*c
- custom instructions
- e.g.,
saturate,ed,multiply-and-accumulate
- e.g.,
- require optimized memories and caches
- custom instructions
- e.g.,
saturate,ed,multiply-and-accumulate
- e.g.,
- require optimized memories and caches
- code may need to be aware of, and explicitly manipulate caches
- custom instructions
- e.g.,
saturate,ed,multiply-and-accumulate
- e.g.,
- require optimized memories and caches
- code may need to be aware of, and explicitly manipulate caches
- may have rudimentary OS but no virtual memory
- general-purpose chips (as opposed to microcontrollers and DSPs)
- general-purpose chips (as opposed to microcontrollers and DSPs)
- more heavy duty computing/memory
- general-purpose chips (as opposed to microcontrollers and DSPs)
- more heavy duty computing/memory
- more flexibility → programming and management of the system
- general-purpose chips (as opposed to microcontrollers and DSPs)
- more heavy duty computing/memory
- more flexibility → programming and management of the system
- commodity processor architectures (e.g,, ARM, Intel x86)
| component | details |
|---|---|
| cores | single or multicore; powerful |
| component | details |
|---|---|
| cores | single or multicore; powerful |
| pipelines | more complex pipelines; better performance, harder to analyze (e.g., wcet) |
| component | details |
|---|---|
| cores | single or multicore; powerful |
| pipelines | more complex pipelines; better performance, harder to analyze (e.g., wcet) |
| clock speeds | higher clock speeds; 100s of khz, or even GHz |
| component | details |
|---|---|
| cores | single or multicore; powerful |
| pipelines | more complex pipelines; better performance, harder to analyze (e.g., wcet) |
| clock speeds | higher clock speeds; 100s of khz, or even GHz |
| ISA | common ISA; well understood, not custom |
| component | details |
|---|---|
| memory | significant memory; megabytes, even gigabytes |
| component | details |
|---|---|
| memory | significant memory; megabytes, even gigabytes |
| cache hierarchies | multiple levels, optimized |
| component | details |
|---|---|
| memory | significant memory; megabytes, even gigabytes |
| cache hierarchies | multiple levels, optimized |
| power consumption | much higher, but can be reduced (e.g., via voltage and frequency scaling) |
| component | details |
|---|---|
| memory | significant memory; megabytes, even gigabytes |
| cache hierarchies | multiple levels, optimized |
| power consumption | much higher, but can be reduced (e.g., via voltage and frequency scaling) |
| size, cost | often higher |
| component | details |
|---|---|
| interrupts, timers | more varied, easily programmable |
| component | details |
|---|---|
| interrupts, timers | more varied, easily programmable |
| I/O | more interfaces, including commodity ones like USB |
| component | details |
|---|---|
| interrupts, timers | more varied, easily programmable |
| I/O | more interfaces, including commodity ones like USB |
| security | often includes additional hardware security features, e.g., ARM TrustZone. |
ARM M-85
most microprocessors do not include → DSPs, ADCs, DACs, etc.
most microprocessors do not include → DSPs, ADCs, DACs, etc.
possible to augment microprocessor → connecting microcontrollers!
- general purpose OS (e.g., Linux, Android, Windows, UNIX, etc.)
- general purpose OS (e.g., Linux, Android, Windows, UNIX, etc.)
- most programming languages and infrastructures (even Docker!)
- general purpose OS (e.g., Linux, Android, Windows, UNIX, etc.)
- most programming languages and infrastructures (even Docker!)
- large number of tooling, analysis, debugging capabilities
- general purpose OS (e.g., Linux, Android, Windows, UNIX, etc.)
- most programming languages and infrastructures (even Docker!)
- large number of tooling, analysis, debugging capabilities
- complex code can run, but increases analysis difficulty
integrates most components in/around a processor
integrates most components in/around a processor → single circuit
integrates most components in/around a processor → single circuit
- processor/chip (microcontroller/microprocessor)
integrates most components in/around a processor → single circuit
- processor/chip (microcontroller/microprocessor)
- memory and memory interfaces
integrates most components in/around a processor → single circuit
- processor/chip (microcontroller/microprocessor)
- memory and memory interfaces
- I/O devices
integrates most components in/around a processor → single circuit
- processor/chip (microcontroller/microprocessor)
- memory and memory interfaces
- I/O devices
- buses → memory and I/O
integrates most components in/around a processor → single circuit
- processor/chip (microcontroller/microprocessor)
- memory and memory interfaces
- I/O devices
- buses → memory and I/O
- storage (e.g., flash) or secondary storage
integrates most components in/around a processor → single circuit
- processor/chip (microcontroller/microprocessor)
- memory and memory interfaces
- I/O devices
- buses → memory and I/O
- storage (e.g., flash) or secondary storage
- radio modems
integrates most components in/around a processor → single circuit
- processor/chip (microcontroller/microprocessor)
- memory and memory interfaces
- I/O devices
- buses → memory and I/O
- storage (e.g., flash) or secondary storage
- radio modems
- (sometimes) accelerators such as GPUs
all placed on a single substrate
- often designed in
C++,MATLAB,SystemC, etc.
- often designed in
C++,MATLAB,SystemC, etc. - once the hardware architectures defined,
- often designed in
C++,MATLAB,SystemC, etc. - once the hardware architectures defined,
- additional elements → hardware description languages
- e.g., register transfer levels (
RTL)
additional components could include,
- DAC
- ADC
additional components could include,
- DAC
- ADC
- radio/signal processing
- wireless modems
additional components could include,
- DAC
- ADC
- radio/signal processing
- wireless modems
- programmable logic
additional components could include,
- DAC
- ADC
- radio/signal processing
- wireless modems
- programmable logic
- networks on chip (NoC)
examples of modern SoCs:
| effect | benefit | problems |
|---|---|---|
| tight integration | better performance, fewer latencies | cannot replace individual components |
| effect | benefit | problems |
|---|---|---|
| tight integration | better performance, fewer latencies | cannot replace individual components |
| custom code/firmware | better use of hardware | not reusable in other systems |
| effect | benefit | problems |
|---|---|---|
| tight integration | better performance, fewer latencies | cannot replace individual components |
| custom code/firmware | better use of hardware | not reusable in other systems |
| custom sw libraries | easier programming | reduces code reusability in other systems |
| effect | benefit | problems |
|---|---|---|
| tight integration | better performance, fewer latencies | cannot replace individual components |
| custom code/firmware | better use of hardware | not reusable in other systems |
| custom sw libraries | easier programming | reduces code reusability in other systems |
| low power consumption | better battery/heat | (potentially) slower |
can vary → depends on SoC
can vary → depends on SoC
- contemporary operating systems → optimized for SoC
- e.g., Raspbian (a Linux variant)
can vary → depends on SoC
- contemporary operating systems → optimized for SoC
- e.g., Raspbian (a Linux variant)
- common programming languages
- e.g.,
C,C++,python,java, evenlisp!
- e.g.,
additional circuitry to speed up certain functions
additional circuitry to speed up certain functions
- idea → certain computation can be offloaded
additional circuitry to speed up certain functions
- idea → certain computation can be offloaded
- main CPU continues to process other code/requests
specialized for certain functions, e.g.,
- parallel matrix multiplications on GPUs
- AES encryption
- some chips include FPGA fabric → custom accelerators
NVidia Jetson Orin → AI/GPU focussed processor
- 1300 MHz clock speeds
- 64 GB Memory
- 256 bit memory bus
- 204 GB/s bandwidth
- supports a variety of graphics features
- DirectX, OpenGL, OpenCL, CUDA, Vulkan, Shader Models
- maximum of 60W power
- 275 trillion operations/s (TOPS)!
- Application-specific integrated circuits (ASICs) and
- field programmable gate arrays (FPGAs)
- semiconductor devices with programmable logic gates
| asic | fpga |
|---|---|
![]() |
![]() |
read more from online text














