System-on-Chip (SoC) Design

ECE382M.20, Fall 2026


Lab #1

Due: 6:00pm, September 22, 2026

 

Instructions:

•        This lab is a team exercise (teams of 2-3 students).

•        Please use the discussion board on ed for Q&A.

•        Submit the report on Canvas.

•        Please check relevant web pages.

 


1         Overview

The goals of this lab are to:

•        Learn the structure of the Llama.cpp source code and compile the code for the Arm platform

•        Identify and propose ways to remove bottlenecks when running on the Arm platform 

The assignment of this lab includes the following:

•        Set up the design and board environment

•        Use a profiling tool to identify the time-consuming portions of the code

•        Benchmark the accuracy and speed of different large language models (LLM)

•        Explore how different optimizations affect inference speed

•        Isolate modules of Llama.cpp and understand how different quantized models pack values

 


2         Environment Setup

Lab work for this class will use the ECE LRC servers and your Ultra96 board. For Lab 1, you can compile the application either on the board or cross-compile it on the LRC servers, but our target is the ARM platform, i.e., you should run all benchmarks and profiling on the board.

 

a)     ECE Linux Servers

We will be using the LRC servers for the class. Instructions for remote access via ssh are listed here: https://wikis.utexas.edu/display/eceit/ECE+Linux+Application+Servers

You can use the /misc/scratch directory on the LRC machines as your own workspace. The scratch directory will not be wiped out until the end of the semester. However, scratch space is also not backed up, i.e. use at your own risk. Execute the following commands:

% cd /misc/scratch

% mkdir <your username>

For software development targeting the board, we will be using Xilinx’s SDK that matches the Ubuntu setup (GCC compiler version) on the board. The SDK includes the capability to compile and link applications for the board using the aarch64-linux-gnu-gcc cross-compiler tool chain, which is installed on the LRC machines and provided by Xilinx together with their development environment.

% module load xilinx/2022
% source /usr/local/packages/Xilinx_2022.2/Vivado/2022.2/settings64.sh

During cross-compilation, Llama.cpp still compiles a few files natively to aid with the build process. To prevent compilation errors, we also need to load a more modern version of GCC.

% module load gcc/14

b)     Boards

Each team will get an Ultra96 board pre-installed with a modified version of Ubuntu 22.04 (PYNQ). You can connect to the board initially from a Linux or Windows host via USB-UART as follows:

1.      Power on and connect the Ultra96 board to the host machine using the provided USB-UART serial cable. If the board doesn’t boot automatically, press the Power Button (SW4). The blue Power On and Done LEDs (D1/D2) next to the microSD card socket should be on.

2.      On a Linux host, search the kernel messaging with the command dmesg|grep tty and look for an indication that the USB-UART is enumerated as a device (typically listed as /dev/ttyUSB1). Connect the device with the minicom application, using the following command: 

% minicom –D /dev/ttyUSB1 –b 115200 -8 -o

The minicom terminal will connect and allow the Ultra96 board terminal output to be interacted with.

3.      On Windows, go into the Device Manager to find the COM port for the USB connection and use a terminal application e.g., PuTTY to connect with a baudrate of 115200.

For further details about the board and its bring-up, you can consult this guide. See this getting started link for general details about setting up the serial connection. If the device driver for the USB UART is not automatically installed, or for further troubleshooting, please see the USB-to-JTAG/UART pod documentation by Avnet.

The login/password will be provided with the board. This account has root access via sudo. To setup Wifi on the board, first put the SSID and pre-shared key (PSK) of the network to connect into in the /root/wpa_supplicant.conf file. To generate the PSK from a plain-text password, run:

% wpa_passphrase <ssid> <password>

and copy and paste the PSK entry into /root/wpa_supplicant.conf.

If you are on campus, you need to use the “utexas-iot” network. The boards share a common PSK value for the utexas IoT network, which can be found in:

/root/wpa_supplicant.conf.Utexas-IOT

If you can’t find this file in your filesystem, please contact the TA.

Then start Wifi with:

% sudo /root/wifi.sh

This command may take ~30s to execute, but as long as the SSID and PSK are correct, it should connect. Sometimes the Wifi on the board can be a little finnicky to connect and may require running the script twice. To run an SSH server on the board, you can follow this guide. You can then connect your board to the network and potentially use ssh to access the board remotely via Wifi. Install any necessary tools/libraries as you wish.

Important: Before unplugging the board from a power source, always make sure to first run:

% sudo halt

It will take a few seconds for the kernel to halt. You can then unplug the board safely.

 


3         Cloning and Compiling the llama.cpp Source Code

You can compile the application directly on the board (slow) or cross-compile it on the LRC servers (much faster):

a)     Get the latest Llama.cpp code from the following link:

% git clone https://github.com/ggml-org/llama.cpp

b)     Go to the new Llama.cpp directory that should have been created:

% cd llama.cpp

c)      Compile the llama.cpp sources. We strongly recommend cross-compiling on the LRC servers as compilation on the Ultra96 board takes approximately 2.5 hours. To cross-compile llama.cpp, we provide a toolchain file which you can copy into your llama.cpp directory.

To build llama.cpp in a new directory run the commands:

% cp /home/projects/gerstl/ece382m/ultra96_toolchain.cmake .
% mkdir build && cd build
% /usr/bin/cmake -DCMAKE_TOOLCHAIN_FILE=../ultra96_toolchain.cmake -DBUILD_SHARED_LIBS=OFF ..
% make -j 8

d)     Compilation should take a few minutes. Forcing static libraries makes it easier to copy the binaries across to the Ultra96 board without dependencies. To sanity check the build, we will use a tiny LLM that generates continuous (mostly nonsensical) prose with the vocabulary of a 3-year-old. Later in this lab we will use larger and more capable models.

On the board run the following command to fetch the Tiny-stories model from Hugging Face:

% wget https://huggingface.co/afrideva/Tinystories-gpt-0.1-3m-GGUF/resolve/main/tinystories-gpt-0.1-3m.Q8_0.gguf

e)     If you cross-compiled on the LRC machines, transfer the binary llama-cli to the board. To test Llama.cpp, run the following command on the board:

% ./llama-cli -m tinystories-gpt-0.1-3m.Q8_0.gguf

The tiny-stories model should take under a minute to load before prompting you to input. This LLM will not respond in a meaningful way but instead generates continuous prose in the style of a children’s story. When the LLM reaches its token limit, it will return control to you and print its performance in tokens/s.

For more information on LLMs and Llama.cpp, you can read the material provided at the following links:

·        https://jalammar.github.io/illustrated-transformer/ (transformer architectures)

·        https://huggingface.co/blog/introduction-to-ggml (the framework kernel)

·        https://huggingface.co/docs/hub/en/gguf (the file format used for open models)

 


4         Benchmarking LLMs

Now that we have built llama.cpp and deployed it on the Ultra96 board, we will use it to run and compare different LLMs. Typically, when choosing an LLM for a given application, we have to trade-off computation requirements (size/speed) against task accuracy. To help assess these trade-offs and compare LLMs, we use benchmarks. In this section, we will install different LLMs and compare them with a common benchmark script.

a)     First, download the Qwen2.5-0.5B parameter instruct model, which is far more capable than Tiny-Stories but will fit within 2 GB of memory available on your Ultra96 platform. This will be the baseline model for the remainder of the lab.

% wget https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct-GGUF/resolve/main/qwen2.5-0.5b-instruct-q8_0.gguf

This model is significantly larger (676 MB) than Tinystories and may take several minutes to download to the board. Try prompting Qwen2.5-0.5B and note its response quality and speed compared to Tinystories.

b)     Qwen-2.5-0.5B is only one of many open-source LLMs. On Hugging Face or otherwise, choose at least two more LLMs and run them using Llama.cpp on your Ultra96 board. Comment on the different models you have chosen, their parameter counts, and network architectures (e.g., using the Netron app to inspect). Note that the maximum model size will be constrained by the 2GB of memory available on the Ultra96 board.

c)      To compare different models and implementations, we will use a custom Python benchmark script which executes a small set of tasks in Llama.cpp and reports speed and task accuracy. Tasks are sampled from three different commonly used datasets: ARC-easy, Hellaswag, and IFEval. The benchmark records the speed (token/s) and accuracy (% score) across all tasks. Investigate the datasets using the provided links. How will we know if our LLM has responded “correctly” for different tasks?

d)     To run the benchmarks, copy and unzip /home/projects/gerstl/ece382m/llama-bench.tar.gz and upload bin/llama-server to the Ultra96 board. To execute the benchmark suite, run:

% python benchmark.py --llama-server <your llama-server path> --model <your model path>

You should see regular prints as each task runs and is evaluated. Report the final speed and scores for Qwen2.5-0.5B.

e)     Quantization is an important optimization used to reduce memory usage in LLMs. If you browse Hugging-Face you will notice that models are available at a range of quantization levels, which allows you to trade off memory usage against task performance. Report different common quantization strategies (e.g., types used, common bit-widths) and describe how values are compressed and unpacked. Investigate how Llama.cpp represents values (weights, KV caches, and activations). At what precision are intermediate values stored and how does this impact accuracy? Using the benchmark script, report speed and test scores for Qwen 2.5-0.5B at Q8_0 and two other quantization levels. Models (GGUF) should be downloaded from: https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct-GGUF

f)       Llama.cpp includes highly optimized code for different CPU/GPU/other hardware architectures. Identify two or more different optimization features/options available on the Ultra96 (AArch64) platform. Profile how these features impact performance (running Qwen2.5-0.5B using Q8_0 quantization).

 


5         Profiling llama.cpp

Next you will identify the performance bottlenecks in the Llama.cpp code and report on your results. Due to tight power constraints in the final project, the target platform is a single A53 CPU at 300 MHz. In this section, we modify the CPU configuration to emulate this target environment and use it to profile llama.cpp.

a)     To configure the board, turn off CPUs 1-3 and lower the frequency of CPU 0 to ~300 MHz:

% echo 299999 | sudo tee /sys/devices/system/cpu/cpu0/cpufreq/scaling_setspeed
% sudo chcpu -d 1-3

 If you want to revert the core frequency back to 1.2 GHz and enable all cores, run:

% echo 1199999 | sudo tee /sys/devices/system/cpu/cpu0/cpufreq/scaling_setspeed
% sudo chcpu -e 0-3

Note these settings reset whenever you reboot the board.

b)     Before you can profile your program, you must first recompile it for profiling. CMake can be finnicky especially in how it caches its build configurations – the safest approach is often to delete the build directory and start again from scratch. Then, add the -pg option to CFLAGS and CXXFLAGS, recompile the code and copy the new llama-cli binary to the board.

% /usr/bin/cmake -DCMAKE_TOOLCHAIN_FILE=../ultra96_toolchain.cmake -DBUILD_SHARED_LIBS=OFF -DCMAKE_C_FLAGS="-pg" -DCMAKE_CXX_FLAGS="-pg" ..
% make -j 8

c)      Now profile the code running single-threaded on the ultra96 board with a short prompt:

% ./llama-cli -m qwen2.5-0.5b-instruct-q8_0.gguf -t 1 -tb 1 -st -p "Describe a system-on-chip in two sentences."

d)     This will take several minutes to finish. Running the program to completion causes a file named gmon.out to be created in the current directory. The gprof tool works by analyzing the data collected during the execution of your program after your program has finished running. Then, gmon.out holds this data in a gprof-readable format.

e)     Run gprof as follows:

% gprof llama-cli gmon.out > llama.perf

f)       Identify the bottleneck(s) of the code based on the execution time of each function. Report your profiling results.

 


Lab Report Submission

This submission should be a report which you submit on Canvas.

The report should contain speed and accuracy results for Qwen2.5-0.5B and any other models you have chosen to compare. The report should explain how Llama.cpp handles quantization within its computation graph and describe different hardware optimizations within the GGML framework. The report should contain profiling results for Qwen2.5-0.5B at 8-bit quantization and should list the performance bottlenecks identified during profiling.