Skip to content

Latest commit

 

History

History
1568 lines (1031 loc) · 34.8 KB

File metadata and controls

1568 lines (1031 loc) · 34.8 KB
title Embedded Architectures
textbook #embedded-architectures

embedded architectures

Design of Autonomous Systems

csci 6907/4907-Section 86

Prof. Sibin Mohan


what is an "embedded system"?


what is an "embedded system"?

  • like autonomy, no exact definition

what is an "embedded system"?

  • like autonomy, no exact definition
  • systems created for specific functionality

what is an "embedded system"?

  • like autonomy, no exact definition
  • systems created for specific functionality
  • fixed/operational for years, decades even!

tradeoff between



performance vs


tradeoff between



performance vs

  • power/battery life

tradeoff between



performance vs

  • power/battery life
  • less memory

tradeoff between



performance vs

  • power/battery life
  • less memory
  • fewer peripherals

tradeoff between



performance vs

  • power/battery life
  • less memory
  • fewer peripherals
  • limited applications

tradeoff between



performance vs

  • power/battery life
  • less memory
  • fewer peripherals
  • limited applications
  • smaller operating systems
  • etc.

main reason?

predictability


predictability

guarantee the system works,


predictability

guarantee the system works,

  • correctly
  • safely

predictability

guarantee the system works,

  • correctly
  • safely

essentially → must be easy to certify the system


the wcet problem


the wcet problem

"worst-case execution time"


the wcet problem

"worst-case execution time"

longest execution time for a program


the wcet problem

"worst-case execution time"

longest execution time for a program
on a specific hardware platform


the wcet problem

"worst-case execution time"

longest execution time for a program
on a specific hardware platform

must consider → all possible inputs


wcet | used to ensure...

  • "schedulability"
  • resource requirements
  • performance limits

of embedded and real-time programs


wcet approaches

lots of approaches to computing the WCET, e.g.,

  • dynamic/empirical → run program lots of times on the platform
  • static → compile time analysis to compute the worst-case paths
  • hybrid → a combination of the two
  • probabilistic → combine dynamic analysis+statistical methods
  • ML-based methods → applying machine-learning to the problem

so what's the "problem"?


so what's the "problem"?

at a high-level, the execution time distributions of applications:

Note: we see that the various methods either underestimate or overestimate the wcet. the former affects the safety, the latter wastes resources


wcet analysis is a hard problem!

structural problems (in hardware/software) make it more difficult!


consider this example

void main()
{
    int max = 10 ;
    int sum = 0;
    for( int i = 0 ; i < max ; ++i)
        sum += i ;
}

consider this example

void main()
{
    int max = 10 ;
    int sum = 0;
    for( int i = 0 ; i < max ; ++i)
        sum += i ;
}

how do you compute the WCET for this code?


need additional information


need additional information

first, define the processor it runs on → say, "P"


need additional information

first, define the processor it runs on → say, "P"

  • how long each instruction takes to execute on P

need additional information

first, define the processor it runs on → say, "P"

  • how long each instruction takes to execute on P
  • how many loop iterations?

need additional information

first, define the processor it runs on → say, "P"

  • how long each instruction takes to execute on P
  • how many loop iterations?
  • what are the startup/cleanup times for the program on P?

adding in all that information, we have...

void main()         // startup cost = 100 cycles
{
int max = 15 ;  // 10 cycles
int sum = 0;    // 10 cycles 
for( int i = 0 ; i < max ; ++i) // 5 cycles, once
sum += i ; // 20 cycles each iteration
}                   // cleanup cost = 120 cycles

to calculate the wcet

1   void main()         // startup cost = 100 cycles
2   {
3       int max = 15 ;  // 10 cycles
4       int sum = 0;    // 10 cycles 
5       for( int i = 0 ; i < max ; ++i) // 5 cycles, once
6            sum += i ; // 20 cycles each iteration
7   }                   // cleanup cost = 120 cycles

$$ wcet = line_1 + line_3 + line_4 + line_5 + (line_6 * max) $$


to calculate the wcet

1   void main()         // startup cost = 100 cycles
2   {
3       int max = 15 ;  // 10 cycles
4       int sum = 0;    // 10 cycles 
5       for( int i = 0 ; i < max ; ++i) // 5 cycles, once
6            sum += i ; // 20 cycles each iteration
7   }                   // cleanup cost = 120 cycles

$$ wcet = line_1 + line_3 + line_4 + line_5 + (line_6 * max) $$

which is fine for this simple example


Now consider this slight change to the above code:

void main( int argc, char* argv[] )
{
    int max = atoi( argv[1] ) ;     // convert the command line arg to max
    int sum = 0; 
    for( int i = 0 ; i < max ; ++i) // how many iterations?
        sum += i ;
}

Now consider this slight change to the above code:

void main( int argc, char* argv[] )
{
    int max = atoi( argv[1] ) ;     // convert the command line arg to max
    int sum = 0; 
    for( int i = 0 ; i < max ; ++i) // how many iterations?
        sum += i ;
}

previous equation fails!


Now consider this slight change to the above code:

void main( int argc, char* argv[] )
{
    int max = atoi( argv[1] ) ;     // convert the command line arg to max
    int sum = 0; 
    for( int i = 0 ; i < max ; ++i) // how many iterations?
        sum += i ;
}

previous equation fails!

no longer know the value of max → arbitrary wcet!


let's make another change to the code:

#define VERY_LARGE_ARRAY+SIZE 1>>18

void main()
{
    int first_array[VERY_LARGE_ARRAY_SIZE] ;
    int second_array[VERY_LARGE_ARRAY_SIZE] ;

    int sum_first = 0;
    int sum_second = 0;
    for( int i = 0 ; i < VERY_LARGE_ARRAY_SIZE * 2 ; ++i)
    {
        if( i%2 )
            first_sum += first_array[i/2] ;
        else
            second_sum += second_array[(int)((i/2)+1)] ;
    }
}

let's make another change to the code:

#define VERY_LARGE_ARRAY+SIZE 1>>18

void main()
{
    int first_array[VERY_LARGE_ARRAY_SIZE] ;
    int second_array[VERY_LARGE_ARRAY_SIZE] ;

    int sum_first = 0;
    int sum_second = 0;
    for( int i = 0 ; i < VERY_LARGE_ARRAY_SIZE * 2 ; ++i)
    {
        if( i%2 )
            first_sum += first_array[i/2] ;
        else
            second_sum += second_array[(int)((i/2)+1)] ;
    }
}

can we compute wcet easily (using previous equation)?

note: maximum size of loop is known → VERY_LARGE_ARRAY_SIZE * 2

Note: pause and ask students what, if anything can go wrong?


let's see what happens in each iteration


let's see what happens in each iteration

iter operation cache state reason
1 first_array loaded miss evicts whatever was previously in cache

let's see what happens in each iteration

iter operation cache state reason
1 first_array loaded miss evicts whatever was previously in cache
2 second_array loaded miss evicts first_array due to lack of space

let's see what happens in each iteration

iter operation cache state reason
1 first_array loaded miss evicts whatever was previously in cache
2 second_array loaded miss evicts first_array due to lack of space
3 first_array loaded again miss evicts second_array due to lack of space
...

program will constantly suffer cache misses!


program will constantly suffer cache misses!

loop's execution time blows up!


loop's execution time blows up!

#define VERY_LARGE_ARRAY+SIZE 1>>18

void main()
{
    int first_array[VERY_LARGE_ARRAY_SIZE] ;
    int second_array[VERY_LARGE_ARRAY_SIZE] ;

    int sum_first = 0;
    int sum_second = 0;
    for( int i = 0 ; i < VERY_LARGE_ARRAY_SIZE * 2 ; ++i)
    {
        if( i%2 )
            first_sum += first_array[i/2] ;     
            // 20 cycles + cache miss time (100s of cycles)
        else
            second_sum += second_array[(int)((i/2)+1)] ; 
            // 20 cycles + cache miss time (100s of cycles)
    }
}

program will constantly suffer cache misses!

loop's execution time blows up!


now we need to model → cache behavior for program/data


further complications from

other hardware features, e.g.,

  • processor pipelining
  • prefetching
  • branch prediction
  • multithreading
  • multicore systems
  • memory buses
  • networks-on-chip
  • and too many others to recount here...

processor feature → improves performance → bad for wcet analysis!


processor feature → improves performance → bad for wcet analysis!

embedded/real-time systems prefer simpler processors


embedded processors

come in myriad shapes and sizes:

  1. Microcontrollers
  2. Digital Signal Processors
  3. Microprocessors
  4. System-on-a-Chip
  5. Embedded accelerators
  6. ASICs and FPGAs

so, we focus on:


Microcontrollers


Microcontrollers

from Wikipedia,

"A microcontroller (MC, UC, or μC) or microcontroller unit (MCU) is a small computer
on a single integrated circuit."


Microcontrollers

  • most common type of "processors" in embedded systems
  • more than 55% of the world's processors!

Microcontrollers

  • most common type of "processors" in embedded systems
  • more than 55% of the world's processors!
  • small, yet critical, systems

Microcontrollers

  • most common type of "processors" in embedded systems
  • more than 55% of the world's processors!
  • small, yet critical, systems
    • car engine control
    • implantable medical devices
    • thermal monitoring
    • millions (billions?) of other applications!

microcontroller hardware features


microcontroller hardware features

component details
one CPU core typically simple 4 or 8 bit chips

Note: Processor bit width refers to the amount of data (in bits) a CPU can process in a single operation, determining its data handling capacity and memory addressing ability, with 32-bit and 64-bit being common standards, where 64-bit CPUs handle larger chunks of data (8 bytes) and access vastly more RAM than 32-bit ones (4 bytes)


microcontroller hardware features

component details
one CPU core typically simple 4 or 8 bit chips
small pipelined architectures sometimes 2 or 4 stage pipelines

microcontroller hardware features

component details
one CPU core typically simple 4 or 8 bit chips
small pipelined architectures sometimes 2 or 4 stage pipelines
some limited memory typically a few hundred kilobytes, perhaps in the form of EEPROMs or FLASH

microcontroller hardware features

component details
one CPU core typically simple 4 or 8 bit chips
small pipelined architectures sometimes 2 or 4 stage pipelines
some limited memory typically a few hundred kilobytes, perhaps in the form of EEPROMs or FLASH
programmable I/O to interact with the real world

microcontroller hardware features [contd.]

component details
low operating frequencies e.g., 4 KHz; simpler/older processors, yet more predictable

Note: processor "frequency is the rate at which a CPU executes basic operations, measured in gigahertz (GHz), indicating billions of cycles per second, essentially how many "ticks" the processor makes to process instructions;


microcontroller hardware features [contd.]

component details
low operating frequencies e.g., 4 KHz; simpler/older processors, yet more predictable
low power consumption in the milliwatts or microwatts ranges; might even be nanowatts when the system is sleeping

microcontroller hardware features [contd.]

component details
low operating frequencies e.g., 4 KHz; simpler/older processors, yet more predictable
low power consumption in the milliwatts or microwatts ranges; might even be nanowatts when the system is sleeping
interrupts (some programmable) often real-time (fixed/low latency)

Note: this is the crux of how we can implement "real-time" behavior. Low latency/deterministic interrupts


microcontroller hardware features [contd.]

component details
low operating frequencies e.g., 4 KHz; simpler/older processors, yet more predictable
low power consumption in the milliwatts or microwatts ranges; might even be nanowatts when the system is sleeping
interrupts (some programmable) often real-time (fixed/low latency)
general-purpose I/O (GPIO) pins for I/O

microcontroller hardware features [contd.]

component details
low operating frequencies e.g., 4 KHz; simpler/older processors, yet more predictable
low power consumption in the milliwatts or microwatts ranges; might even be nanowatts when the system is sleeping
interrupts (some programmable) often real-time (fixed/low latency)
several general-purpose I/O (GPIO) pins for I/O
timers e.g., a programmable interval timer (PIT)

Additional features found on some microcontrollers:


Additional features found on some microcontrollers:

component details
analog to digital (ADC) convertors to convert incoming (real-world, sensor) data to a digital form that the uC can operate on

Note:

  • because the real world is "analog"/continuous

Additional features found on some microcontrollers:

component details
analog to digital (ADC) convertors to convert incoming (real-world, sensor) data to a digital form that the uC can operate on
digital-to-analog (DAC) convertor to do the opposite, convert from digital to analog signals to send outputs in that form

Note:

  • because the real world is "analog"/continuous

Additional features found on some microcontrollers:

component details
analog to digital (ADC) convertors to convert incoming (real-world, sensor) data to a digital form that the uC can operate on
digital-to-analog (DAC) convertor to do the opposite, convert from digital to analog signals to send outputs in that form
universal async tx/rx (UART) to receive/send data over a serial line

Additional features [contd.]:

component details
pulse width modulation (PWM) so that the CPU can control motors (significant for us in autonomous/automotive systems), power systems, resistive loads, etc.

Additional features found on some microcontrollers:

component details
pulse width modulation (PWM) so that the CPU can control motors (significant for us in autonomous/automotive systems), power systems, resistive loads, etc.
JTAG interface debugging interface

Microcontroller programs and data


Microcontroller programs and data

  • are small → must fit in (limited) memory

Microcontroller programs and data

  • are small → must fit in (limited) memory
  • often directly programmed in assembly!

Microcontroller programs and data

  • are small → must fit in (limited) memory
  • often directly programmed in assembly!
    • assembly code might need hand tuning

Note: hand-tuning is for both, performance as well as fitting into the limited memory


Microcontroller programs and data

  • are small → must fit in (limited) memory
  • often directly programmed in assembly!
    • assembly code might need hand tuning
  • C is another popular language

Microcontroller programs and data

  • are small → must fit in (limited) memory
  • often directly programmed in assembly!
    • assembly code might need hand tuning
  • C is another popular language
  • no operating systems (or very rare)!

Note:

  • "**bare metal execution"

Microcontroller programs and data

  • are small → must fit in (limited) memory
  • often directly programmed in assembly!
    • assembly code might need hand tuning
  • C is another popular language
  • no operating systems (or very rare)!
  • sometimes use special-purpose programming languages/instructions

popular microcontroller families:


Atmel ATmega

Microchip Technology

Motorola (Freescale)

NXP

digital signal processor (dsp)


what is a dsp?

  • specialized microcontrollers optimized for digital signal processing

what is a dsp?

  • specialized microcontrollers optimized for digital signal processing


what is a dsp?

  • specialized microcontrollers optimized for digital signal processing
  • widely used in:
    • audio processing
    • radar and sonar
    • speech recognition systems
    • image processing
    • satellites
    • telecommunications
    • mobile phones
    • televisions

main goals of dsp

  • isolate, measure, compress, filter analog signals in the real world
  • often have stringent real-time constraints

example

  • texas instruments dsp chip, tms320 series
  • one of the most famous examples of dsp systems


typical digital signal processing

  • repetitive mathematical operations
    • over a large number of samples
    • in real-time

typical digital signal processing

  • repetitive mathematical operations
    • over a large number of samples
    • in real-time


"multiply-accumulate" (mac) algorithm

Note: Imagine you are reading a sensor that measures temperature every second. The signal is "noisy" (it jumps around erratically). To smooth it, you don't just look at the current number; you look at the average of the current number and the previous two numbers.

This requires two repetitive steps that are the heart of DSP:

Multiplication: Weighting the data points.

Accumulation: Adding them together.


"multiply-accumulate" (mac) algorithm

signal = [10, 22, 10, 12]
weights = [0.5, 0.25, 0.25]
output = []

# This loop is the "repetitive" nature of DSP
for i in range(2, len(signal)):
    # The MAC (Multiply-Accumulate) Operation
    val = (signal[i]   * weights[0]) + \
          (signal[i-1] * weights[1]) + \
          (signal[i-2] * weights[2])
    output.append(val)

print(output) # Result: [13.5, 14.0]

steps involved:

  • analog to digital conversion
  • manipulation (the core algorithm)
  • digital to analog conversion

real-time constraints

  • entire process → completed with low latency, fixed deadline
  • low power requirements → battery-constrained devices

software implementations exist

Note: hardware is faster, more predictable, low latency/jitter


Typical DSP architecture


additional details


additional details

  • custom instructions
    • e.g., saturate, ed, multiply-and-accumulate

Note:

  • caps the minimum or maximum value that can be held in a fixed-point representation (e.g., moving audio knob past 100 → shouldn't "wrap around" to 0),
  • Euclidean distance, a+b*c

additional details

  • custom instructions
    • e.g., saturate, ed, multiply-and-accumulate
  • require optimized memories and caches

additional details

  • custom instructions
    • e.g., saturate, ed, multiply-and-accumulate
  • require optimized memories and caches
  • code may need to be aware of, and explicitly manipulate caches

additional details

  • custom instructions
    • e.g., saturate, ed, multiply-and-accumulate
  • require optimized memories and caches
  • code may need to be aware of, and explicitly manipulate caches
  • may have rudimentary OS but no virtual memory

Microprocessors


Microprocessors

  • general-purpose chips (as opposed to microcontrollers and DSPs)

Microprocessors

  • general-purpose chips (as opposed to microcontrollers and DSPs)
  • more heavy duty computing/memory

Microprocessors

  • general-purpose chips (as opposed to microcontrollers and DSPs)
  • more heavy duty computing/memory
  • more flexibility → programming and management of the system

Microprocessors

  • general-purpose chips (as opposed to microcontrollers and DSPs)
  • more heavy duty computing/memory
  • more flexibility → programming and management of the system
  • commodity processor architectures (e.g,, ARM, Intel x86)

Main features of microprocessors


Main features of microprocessors

component details
cores single or multicore; powerful

Main features of microprocessors

component details
cores single or multicore; powerful
pipelines more complex pipelines; better performance, harder to analyze (e.g., wcet)

Main features of microprocessors

component details
cores single or multicore; powerful
pipelines more complex pipelines; better performance, harder to analyze (e.g., wcet)
clock speeds higher clock speeds; 100s of khz, or even GHz

Main features of microprocessors

component details
cores single or multicore; powerful
pipelines more complex pipelines; better performance, harder to analyze (e.g., wcet)
clock speeds higher clock speeds; 100s of khz, or even GHz
ISA common ISA; well understood, not custom

Main features [contd.]

component details
memory significant memory; megabytes, even gigabytes

Main features [contd.]

component details
memory significant memory; megabytes, even gigabytes
cache hierarchies multiple levels, optimized

Main features [contd.]

component details
memory significant memory; megabytes, even gigabytes
cache hierarchies multiple levels, optimized
power consumption much higher, but can be reduced (e.g., via voltage and frequency scaling)

Main features [contd.]

component details
memory significant memory; megabytes, even gigabytes
cache hierarchies multiple levels, optimized
power consumption much higher, but can be reduced (e.g., via voltage and frequency scaling)
size, cost often higher

Main features [contd.]

component details
interrupts, timers more varied, easily programmable

Main features [contd.]

component details
interrupts, timers more varied, easily programmable
I/O more interfaces, including commodity ones like USB

Main features [contd.]

component details
interrupts, timers more varied, easily programmable
I/O more interfaces, including commodity ones like USB
security often includes additional hardware security features, e.g., ARM TrustZone.

ARM M-85


most microprocessors do not include → DSPs, ADCs, DACs, etc.


most microprocessors do not include → DSPs, ADCs, DACs, etc.

possible to augment microprocessor → connecting microcontrollers!


software → most flexibility


software

  • general purpose OS (e.g., Linux, Android, Windows, UNIX, etc.)

software

  • general purpose OS (e.g., Linux, Android, Windows, UNIX, etc.)
  • most programming languages and infrastructures (even Docker!)

software

  • general purpose OS (e.g., Linux, Android, Windows, UNIX, etc.)
  • most programming languages and infrastructures (even Docker!)
  • large number of tooling, analysis, debugging capabilities

software

  • general purpose OS (e.g., Linux, Android, Windows, UNIX, etc.)
  • most programming languages and infrastructures (even Docker!)
  • large number of tooling, analysis, debugging capabilities
  • complex code can run, but increases analysis difficulty

System-on-a-Chip (SoC)


System-on-a-Chip (SoC)

integrates most components in/around a processor


System-on-a-Chip (SoC)

integrates most components in/around a processor → single circuit


System-on-a-Chip (SoC)

integrates most components in/around a processor → single circuit

  • processor/chip (microcontroller/microprocessor)

System-on-a-Chip (SoC)

integrates most components in/around a processor → single circuit

  • processor/chip (microcontroller/microprocessor)
  • memory and memory interfaces

System-on-a-Chip (SoC)

integrates most components in/around a processor → single circuit

  • processor/chip (microcontroller/microprocessor)
  • memory and memory interfaces
  • I/O devices

System-on-a-Chip (SoC)

integrates most components in/around a processor → single circuit

  • processor/chip (microcontroller/microprocessor)
  • memory and memory interfaces
  • I/O devices
  • buses → memory and I/O

System-on-a-Chip (SoC)

integrates most components in/around a processor → single circuit

  • processor/chip (microcontroller/microprocessor)
  • memory and memory interfaces
  • I/O devices
  • buses → memory and I/O
  • storage (e.g., flash) or secondary storage

System-on-a-Chip (SoC)

integrates most components in/around a processor → single circuit

  • processor/chip (microcontroller/microprocessor)
  • memory and memory interfaces
  • I/O devices
  • buses → memory and I/O
  • storage (e.g., flash) or secondary storage
  • radio modems

System-on-a-Chip (SoC)

integrates most components in/around a processor → single circuit

  • processor/chip (microcontroller/microprocessor)
  • memory and memory interfaces
  • I/O devices
  • buses → memory and I/O
  • storage (e.g., flash) or secondary storage
  • radio modems
  • (sometimes) accelerators such as GPUs

all placed on a single substrate


soc design


soc design

  • often designed in C++, MATLAB, SystemC, etc.

soc design

  • often designed in C++, MATLAB, SystemC, etc.
  • once the hardware architectures defined,

soc design

  • often designed in C++, MATLAB, SystemC, etc.
  • once the hardware architectures defined,
    • additional elements → hardware description languages
    • e.g., register transfer levels (RTL)

additional components could include,

  • DAC
  • ADC

additional components could include,

  • DAC
  • ADC
  • radio/signal processing
  • wireless modems

additional components could include,

  • DAC
  • ADC
  • radio/signal processing
  • wireless modems
  • programmable logic

additional components could include,

  • DAC
  • ADC
  • radio/signal processing
  • wireless modems
  • programmable logic
  • networks on chip (NoC)

examples of modern SoCs:

Broadcom Soc from Raspberry Pi

Apple M1 SoC


side-effects of hardware integration


side-effects of hardware integration

effect benefit problems
tight integration better performance, fewer latencies cannot replace individual components

side-effects of hardware integration

effect benefit problems
tight integration better performance, fewer latencies cannot replace individual components
custom code/firmware better use of hardware not reusable in other systems

side-effects of hardware integration

effect benefit problems
tight integration better performance, fewer latencies cannot replace individual components
custom code/firmware better use of hardware not reusable in other systems
custom sw libraries easier programming reduces code reusability in other systems

side-effects of hardware integration

effect benefit problems
tight integration better performance, fewer latencies cannot replace individual components
custom code/firmware better use of hardware not reusable in other systems
custom sw libraries easier programming reduces code reusability in other systems
low power consumption better battery/heat (potentially) slower

software stack


software stack

can vary → depends on SoC


software stack

can vary → depends on SoC

  • contemporary operating systems → optimized for SoC
    • e.g., Raspbian (a Linux variant)

software stack

can vary → depends on SoC

  • contemporary operating systems → optimized for SoC
    • e.g., Raspbian (a Linux variant)
  • common programming languages
    • e.g., C, C++, python, java, even lisp!

embedded accelerators


embedded accelerators

additional circuitry to speed up certain functions


embedded accelerators

additional circuitry to speed up certain functions

  • idea → certain computation can be offloaded

embedded accelerators

additional circuitry to speed up certain functions

  • idea → certain computation can be offloaded
  • main CPU continues to process other code/requests

embedded accelerators | examples


embedded accelerators

specialized for certain functions, e.g.,

  • parallel matrix multiplications on GPUs
  • AES encryption
  • some chips include FPGA fabric → custom accelerators

embedded accelerators

NVidia Jetson Orin → AI/GPU focussed processor


NVidia Jetson Orin | specs

  • 1300 MHz clock speeds
  • 64 GB Memory
  • 256 bit memory bus
  • 204 GB/s bandwidth
  • supports a variety of graphics features
    • DirectX, OpenGL, OpenCL, CUDA, Vulkan, Shader Models
  • maximum of 60W power
  • 275 trillion operations/s (TOPS)!

asic and fpga


asic and fpga

  • Application-specific integrated circuits (ASICs) and
  • field programmable gate arrays (FPGAs)
  • semiconductor devices with programmable logic gates
asic fpga

asic and fpga

read more from online text