System-on-Chip (SoC) Design
ECE382M.20, Fall 2026
Lab #1
Due: 6:00pm, September 22, 2026
Instructions:
•
This lab is a team exercise (teams of 2-3
students).
•
Please use the discussion board on ed for
Q&A.
•
Submit the report on Canvas.
•
Please check relevant web pages.
The goals of this lab are to:
•
Learn the structure of the Llama.cpp source code
and compile the code for the Arm platform
•
Identify and propose ways to remove bottlenecks
when running on the Arm platform
The assignment of this lab includes the
following:
•
Set up the design and board environment
•
Use a profiling tool to identify the time-consuming
portions of the code
•
Benchmark the accuracy and speed of different large
language models (LLM)
•
Explore how different optimizations affect
inference speed
•
Isolate modules of Llama.cpp and understand how
different quantized models pack values
Lab work for this class will use the ECE LRC servers and your Ultra96
board. For Lab 1, you can compile the application either on the board or
cross-compile it on the LRC servers, but our target is the ARM platform, i.e.,
you should run all benchmarks and profiling on the board.
a)
ECE Linux Servers
We will be using the LRC servers for the class. Instructions for remote
access via ssh are listed here: https://wikis.utexas.edu/display/eceit/ECE+Linux+Application+Servers
You can use the /misc/scratch directory on the
LRC machines as your own workspace. The scratch directory will not be wiped out
until the end of the semester. However, scratch space is also not backed up,
i.e. use at your own risk. Execute the following commands:
% cd /misc/scratch
% mkdir <your username>
For software development targeting the board, we will be using Xilinx’s
SDK that matches the Ubuntu setup (GCC compiler version) on the board. The SDK
includes the capability to compile and link applications for the board using the
aarch64-linux-gnu-gcc cross-compiler tool chain, which is installed on the LRC
machines and provided by Xilinx together with their development environment.
% module load xilinx/2022
% source /usr/local/packages/Xilinx_2022.2/Vivado/2022.2/settings64.sh
During cross-compilation, Llama.cpp still compiles a few files
natively to aid with the build process. To prevent compilation errors, we also
need to load a more modern version of GCC.
% module load gcc/14
b)
Boards
Each team will get an Ultra96 board
pre-installed with a modified version of Ubuntu 22.04 (PYNQ). You can connect
to the board initially from a Linux or Windows host via USB-UART as follows:
1.
Power on and connect the Ultra96 board to the host
machine using the provided USB-UART serial cable. If the board doesn’t boot
automatically, press the Power Button (SW4). The blue Power On
and Done LEDs (D1/D2) next to the microSD card socket should be on.
2.
On a Linux host, search the kernel messaging with
the command dmesg|grep tty and look for an
indication that the USB-UART is enumerated as a device (typically listed as /dev/ttyUSB1). Connect the
device with the minicom application,
using the following command:
% minicom –D /dev/ttyUSB1 –b 115200 -8 -o
The minicom terminal will connect and allow the Ultra96 board terminal
output to be interacted with.
3.
On Windows, go into the Device Manager to find the
COM port for the USB connection and use a terminal application e.g., PuTTY to
connect with a baudrate of 115200.
For further details about the board and its
bring-up, you can consult this guide. See this getting started link for general
details about setting up the serial connection. If the device driver for the
USB UART is not automatically installed, or for further troubleshooting, please
see the USB-to-JTAG/UART
pod documentation by Avnet.
The login/password will be provided with the board. This account has
root access via sudo. To setup
Wifi on the board, first put the SSID and pre-shared
key (PSK) of the network to connect into in the /root/wpa_supplicant.conf file. To generate
the PSK from a plain-text password, run:
%
wpa_passphrase <ssid>
<password>
and copy and paste the PSK entry into /root/wpa_supplicant.conf.
If you are on campus, you need to use the “utexas-iot”
network. The boards share a common PSK value for the utexas
IoT network, which can be found in:
/root/wpa_supplicant.conf.Utexas-IOT
If you can’t find this file in your filesystem, please contact the TA.
Then start Wifi with:
% sudo /root/wifi.sh
This command may take ~30s to execute, but as long as
the SSID and PSK are correct, it should connect. Sometimes the Wifi on the board can be a little finnicky to connect and
may require running the script twice. To run an SSH server on the board, you can follow this
guide. You can then connect your board to the network and potentially use ssh to access the board remotely via Wifi. Install any necessary tools/libraries as you wish.
Important: Before unplugging the board from a power source,
always make sure to first run:
%
sudo halt
It will take a few seconds for the kernel to halt.
You can then unplug the board safely.
You can compile the application
directly on the board (slow) or cross-compile it on the LRC servers (much
faster):
a) Get the latest Llama.cpp code from the following link:
% git clone https://github.com/ggml-org/llama.cpp
b) Go to the new Llama.cpp directory that should have been created:
%
cd llama.cpp
c) Compile the llama.cpp sources. We strongly recommend cross-compiling on the LRC servers as compilation on the Ultra96 board takes approximately 2.5 hours. To cross-compile llama.cpp, we provide a toolchain file which you can copy into your llama.cpp directory.
To build llama.cpp in a new directory run the commands:
% cp /home/projects/gerstl/ece382m/ultra96_toolchain.cmake
.
% mkdir build && cd build
% /usr/bin/cmake
-DCMAKE_TOOLCHAIN_FILE=../ultra96_toolchain.cmake -DBUILD_SHARED_LIBS=OFF ..
% make -j 8
d) Compilation should take a few minutes. Forcing static libraries makes it easier to copy the binaries across to the Ultra96 board without dependencies. To sanity check the build, we will use a tiny LLM that generates continuous (mostly nonsensical) prose with the vocabulary of a 3-year-old. Later in this lab we will use larger and more capable models.
On the board run the following command to fetch the Tiny-stories model from Hugging Face:
% wget
https://huggingface.co/afrideva/Tinystories-gpt-0.1-3m-GGUF/resolve/main/tinystories-gpt-0.1-3m.Q8_0.gguf
e) If you cross-compiled on the LRC machines, transfer the binary llama-cli to the board. To test Llama.cpp, run the following command on the board:
% ./llama-cli -m tinystories-gpt-0.1-3m.Q8_0.gguf
The tiny-stories model should take under a minute to load before prompting you to input. This LLM will not respond in a meaningful way but instead generates continuous prose in the style of a children’s story. When the LLM reaches its token limit, it will return control to you and print its performance in tokens/s.
For more information on LLMs and Llama.cpp, you can read the material provided at the following links:
· https://jalammar.github.io/illustrated-transformer/ (transformer architectures)
· https://huggingface.co/blog/introduction-to-ggml (the framework kernel)
· https://huggingface.co/docs/hub/en/gguf (the file format used for open models)
Now that we have built llama.cpp and deployed it on the Ultra96 board,
we will use it to run and compare different LLMs. Typically, when choosing an
LLM for a given application, we have to trade-off computation requirements
(size/speed) against task accuracy. To help assess these trade-offs and compare
LLMs, we use benchmarks. In this section, we will install different LLMs and
compare them with a common benchmark script.
a)
First, download the Qwen2.5-0.5B parameter instruct
model, which is far more capable than Tiny-Stories but will fit within 2 GB of
memory available on your Ultra96 platform. This will be the baseline model for
the remainder of the lab.
% wget
https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct-GGUF/resolve/main/qwen2.5-0.5b-instruct-q8_0.gguf
This model is significantly larger (676 MB) than Tinystories
and may take several minutes to download to the board. Try prompting
Qwen2.5-0.5B and note its response quality and speed compared to Tinystories.
b)
Qwen-2.5-0.5B is only one of many open-source LLMs.
On Hugging Face or otherwise, choose at least two more LLMs and run them using
Llama.cpp on your Ultra96 board. Comment on the different models you have
chosen, their parameter counts, and network architectures (e.g., using the Netron app to
inspect). Note that the maximum model size will be constrained by the 2GB of
memory available on the Ultra96 board.
c)
To compare different models and implementations, we
will use a custom Python benchmark script which executes a small set of tasks
in Llama.cpp and reports speed and task accuracy. Tasks are sampled from three
different commonly used datasets: ARC-easy,
Hellaswag, and IFEval.
The benchmark records the speed (token/s) and accuracy (% score) across all
tasks. Investigate the datasets using the provided links.
How will we know if our LLM has responded “correctly” for different tasks?
d)
To run the benchmarks, copy and unzip /home/projects/gerstl/ece382m/llama-bench.tar.gz and upload bin/llama-server to the Ultra96
board. To execute the benchmark suite, run:
% python benchmark.py --llama-server <your llama-server
path> --model <your model path>
You should see regular prints as each task runs and is evaluated. Report
the final speed and scores for Qwen2.5-0.5B.
e)
Quantization is an important optimization used to
reduce memory usage in LLMs. If you browse Hugging-Face you will notice that
models are available at a range of quantization levels, which allows you to
trade off memory usage against task performance. Report different common
quantization strategies (e.g., types used, common bit-widths) and describe how values
are compressed and unpacked. Investigate how Llama.cpp represents values
(weights, KV caches, and activations). At what precision are intermediate
values stored and how does this impact accuracy? Using the benchmark script,
report speed and test scores for Qwen 2.5-0.5B at Q8_0 and two other
quantization levels. Models (GGUF) should be downloaded from: https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct-GGUF
f)
Llama.cpp includes highly optimized code for
different CPU/GPU/other hardware architectures. Identify two or more different
optimization features/options available on the Ultra96 (AArch64) platform.
Profile how these features impact performance (running Qwen2.5-0.5B using Q8_0
quantization).
Next you will identify the performance bottlenecks in the
Llama.cpp code and report on your results. Due to tight power constraints in the final project,
the target platform is a single A53 CPU at 300 MHz.
In this section, we modify the CPU configuration to emulate this target
environment and use it to profile llama.cpp.
a)
To configure the board, turn off CPUs 1-3 and lower
the frequency of CPU 0 to ~300 MHz:
% echo 299999 | sudo tee /sys/devices/system/cpu/cpu0/cpufreq/scaling_setspeed
% sudo chcpu -d 1-3
If you want
to revert the core frequency back to 1.2 GHz and enable all cores, run:
% echo 1199999 | sudo tee /sys/devices/system/cpu/cpu0/cpufreq/scaling_setspeed
% sudo chcpu -e 0-3
Note these settings reset whenever you reboot the
board.
b)
Before you can profile your program, you must first
recompile it for profiling. CMake can be finnicky
especially in how it caches its build configurations – the safest approach is
often to delete the build directory and start again from scratch. Then, add the
-pg option to CFLAGS and CXXFLAGS,
recompile
the code and copy the new llama-cli binary to the board.
% /usr/bin/cmake -DCMAKE_TOOLCHAIN_FILE=../ultra96_toolchain.cmake
-DBUILD_SHARED_LIBS=OFF -DCMAKE_C_FLAGS="-pg"
-DCMAKE_CXX_FLAGS="-pg"
..
% make
-j 8
c)
Now profile the code running single-threaded on the
ultra96 board with a short prompt:
%
./llama-cli
-m qwen2.5-0.5b-instruct-q8_0.gguf -t 1 -tb 1 -st -p
"Describe a system-on-chip in two sentences."
d)
This will take several minutes to finish. Running
the program to completion causes a file named gmon.out to be created in
the current directory. The gprof tool works by analyzing the data collected
during the execution of your program after your program has finished running.
Then, gmon.out holds this data in a gprof-readable
format.
e)
Run gprof as follows:
% gprof llama-cli gmon.out > llama.perf
f)
Identify the bottleneck(s) of the code based on the
execution time of each function. Report your profiling results.
This submission should be a report which you submit on Canvas.
The report should contain speed and accuracy
results for Qwen2.5-0.5B and any other models you have chosen to compare. The
report should explain how Llama.cpp handles quantization within its computation
graph and describe different hardware optimizations within the GGML framework.
The report should contain profiling results for Qwen2.5-0.5B at 8-bit quantization and should list the performance bottlenecks
identified during profiling.