Last updated
This guide shows how to design an ESP accelerator in C/C++ and synthesize it with PandA Bambu, an open-source high-level synthesis (HLS) tool. The flow follows the same structure as ESP’s other accelerator flows: generate a skeleton, implement and verify the kernel, install the generated RTL, and integrate the accelerator into an SoC.
dev
branch and is planned for the next ESP release. Commands and options may
change before the release. Coming soon
Platform support: the flow targets AMD/Xilinx FPGA technologies. For the Virtex-7 technology library (
virtex7, used by the VC707 and proFPGA XC7V2000T), ESP selects a Bambu device characterization automatically. For other technologies, setBAMBU_DEVICEas described in Configuration variables.
This guide assumes that you are familiar with the ESP infrastructure and know how to configure, simulate and prototype a single-core SoC.
Prerequisites
- An ESP checkout on the
devbranch, set up as described in the setup guide. - PandA Bambu built from the
feature/unified_ac_channelsbranch. ESP validation used revision4150ff6e20d543eb5d241df7085046b6a355f283, Clang 16 and Verilator 4.038. - For Bambu RTL co-simulation, the selected simulator (Verilator by default)
and 32-bit C++ linking support, such as the
g++-multilibpackage on Ubuntu. - (optional) Icarus Verilog (
iverilogandvvp) for the standalone DMA test of the example accelerator.
Point ESP at the Bambu installation. BAMBU_ENV is optional and names a
settings script that ESP sources before each Bambu invocation:
export BAMBU=<bambu_path>/bin/bambu
export BAMBU_ENV=<bambu_path>/settings.sh
If bambu is already on your PATH and needs no settings script, you can
omit both variables.
Generate the accelerator skeleton
Run ESP’s interactive accelerator generator from the ESP root and select B for Bambu HLS:
cd <esp>
./tools/accgen/accgen.sh
=== Initializing ESP accelerator template ===
* Enter accelerator name [dummy]: example
* Select design flow (Stratus HLS, Vivado HLS, hls4ml, Catapult HLS, RTL, Bambu HLS) [S]: B
* Enter ESP path [<esp>]:
* Enter unique accelerator id as three hex digits [04A]: 07B
* Enter accelerator registers
- register 0 name [size]: len
- register 0 default value [1]: 64
- register 0 max value [64]: 1024
- register 1 name []:
...
The Bambu flow currently supports a subset of the generator options. The generator warns and applies these settings automatically:
| Option | Bambu flow value |
|---|---|
| Data bit-width | 32 bits (32-bit elements over a 64-bit DMA interface) |
| Chunking factor | 1 (no chunking) |
| Batching factor | 1 (no batching) |
The accelerator is created in accelerators/bambu_hls/<name>_bambu/:
<name>_bambu/
├── hw
│ ├── <name>_bambu.xml # accelerator description and registers
│ ├── hls
│ │ ├── Makefile -> ../../../common/hls/Makefile
│ │ └── bambu_args.mk # per-accelerator Bambu flags (optional)
│ ├── src
│ │ ├── <name>_bambu.cpp # C/C++ kernel
│ │ └── <name>_bambu_basic_dma64/
│ │ └── <name>_bambu_basic_dma64.v # ESP socket wrapper
│ └── tb
│ └── <name>_bambu_tb.cpp # C++ testbench
└── sw
├── baremetal # bare-metal test application
└── linux # device driver and user-space application
Implement the accelerator
The generated kernel in hw/src/<name>_bambu.cpp is a dataflow design with
concurrent load, compute and store stages connected by ping-pong buffers. The
placeholder computation squares each input element; replace it with your
accelerator’s function. Each configuration register becomes a scalar argument
of the top function and a conf_info_<register> port of the wrapper.
Keep these conventions:
- The C/C++ top function must be named
<name>_bambu_core. Bambu generates a Verilog module with that name, and the ESP socket wrapper instantiates it. - The wrapper in
hw/src/<name>_bambu_basic_dma64/adapts the generated core’s AXI-Stream channels to the ESP accelerator socket. Only a 64-bit DMA wrapper is provided. - Register handling, interrupts, scatter-gather address translation and DMA over the network on chip are provided by the generated ESP socket, not by the Bambu core.
- Update the testbench in
hw/tb/<name>_bambu_tb.cppso that it checks your outputs and printsPASSon success; the co-simulation target looks for that string.
Add extra Bambu options for one accelerator in hw/hls/bambu_args.mk, for
example:
BAMBU_EXTRA_FLAGS = --no-clean
Synthesize and verify
Run the Bambu targets from the working directory of your SoC. The targets are named after the accelerator directory:
cd <esp>/socs/xilinx-vc707-xc7vx485t
# Native C++ execution of the testbench (no HLS)
make <name>_bambu-exe
# High-level synthesis and installation into ESP
make <name>_bambu-hls
# Bambu RTL co-simulation of the testbench
make <name>_bambu-sim
-execompiles the testbench and kernel with the host C++ compiler, using Bambu’s headers.-hlsruns Bambu inhw/hls-work-<tech>/and installs the accelerator XML, the socket wrapper, the generated core and its memory initialization files intech/<tech>/acc/<name>_bambu/. The log is saved in the SoC working directory’s HLS log folder.-simre-runs Bambu with its generated testbench and simulator, and passes only if the testbench log containsPASS.
Other targets:
| Target | Description |
|---|---|
<name>_bambu-clean |
Remove Bambu outputs from the work directory |
<name>_bambu-distclean |
Also remove the installed RTL from tech/ |
bambuhls_acc |
Run HLS for every Bambu accelerator |
bambuhls_acc-clean, bambuhls_acc-distclean |
Clean every Bambu accelerator |
print-available-acc |
List accelerators, including Bambu accelerators |
Bambu accelerators are not built by the Stratus HLS make acc target; use
bambuhls_acc instead.
Integrate the accelerator into an SoC
After make <name>_bambu-hls, the accelerator appears in the ESP
configuration GUI like any other installed accelerator:
- Run
make esp-xconfigand place the accelerator in a tile with thebasic_dma64implementation. Use a configuration with 64-bit DMA, such as the default Ariane configuration. -
Generate the sockets and build the bare-metal test:
make socketgen make <name>_bambu-baremetal -
Run a full-system RTL simulation with the bare-metal test:
make sim TEST_PROGRAM=$PWD/soft-build/ariane/baremetal/<name>_bambu.exe -
Build the bitstream and test on FPGA with the usual single-core SoC targets:
make vivado-syn make fpga-program make fpga-run
ModelSim and Vivado compile Bambu accelerator libraries automatically and
stage their memory initialization files. When simulating with ModelSim, use the
same version that compiled the cached AMD/Xilinx simulation libraries. For an
existing Vivado project, run make vivado-update to pick up a new accelerator.
Example: sqr_bambu
accelerators/bambu_hls/sqr_bambu/ is a validated example. It squares
unsigned 32-bit words modulo 232 with load, compute and store
stages that run concurrently, using two 16-word banks at each stage boundary.
| Register | Offset | Description |
|---|---|---|
len |
0x40 |
Number of 32-bit words; a positive multiple of 16 |
base_in |
0x44 |
Input offset, in 64-bit DMA beats |
base_out |
0x48 |
Output offset, in 64-bit DMA beats |
Use disjoint input and output regions. The hardware does not reject invalid lengths, so software must enforce them.
The reference SoC has one Ariane processor tile, one memory tile, one I/O tile
and one SQR_BAMBU accelerator tile with the basic_dma64 implementation,
caches disabled and 64-bit DMA. Instead of using the GUI, you can add the tile
to socgen/esp/.esp_config directly:
TILE_1_0 = 2 acc SQR_BAMBU basic_dma64 0 0 sld
Build and validate it from the ESP root:
make -C socs/xilinx-vc707-xc7vx485t sqr_bambu-hls
make -C socs/xilinx-vc707-xc7vx485t sqr_bambu-sim
# Standalone DMA test of the installed RTL (requires Icarus Verilog)
bash accelerators/bambu_hls/sqr_bambu/hw/tb/run_dma_test.sh
make -C socs/xilinx-vc707-xc7vx485t esp-xconfig
make -C socs/xilinx-vc707-xc7vx485t socketgen sqr_bambu-baremetal
make -C socs/xilinx-vc707-xc7vx485t sim \
TEST_PROGRAM="$PWD/socs/xilinx-vc707-xc7vx485t/soft-build/ariane/baremetal/sqr_bambu.exe"
make -C socs/xilinx-vc707-xc7vx485t vivado-syn
The bare-metal application runs three 272-word invocations and checks the results. A passing run prints:
sqr_bambu PASS total_errors=0
ESP’s simulation harness then prints Failure: Program Completed! as its
normal end-of-test assertion. Check the application’s PASS line rather than
that assertion.
Verification note:
sqr_bambu-simsetsBAMBU_SKIP_VERIFICATIONbecause Bambu’s sequential C model cannot represent the concurrent bank handoff. The testbench still checks the RTL outputs independently, and the standalone DMA test covers descriptors, stalls and repeated starts.
The example includes a bare-metal application only; it has no Linux driver or user-space application.
Configuration variables
Set these variables in the environment or on the make command line to
override the defaults:
| Variable | Default | Description |
|---|---|---|
BAMBU |
bambu |
Bambu executable, on PATH or as a full path |
BAMBU_ENV |
(empty) | Settings script sourced before running Bambu |
BAMBU_DEVICE |
xc7vx690t,-3,ffg1930 for virtex7 |
Bambu device as device,speedgrade,package; required for other technologies |
BAMBU_CLOCK_PERIOD |
10 |
Target clock period in ns |
BAMBU_FLAGS |
see accelerators/bambu_hls/common/hls/Makefile |
Base Bambu options, including Clang 16 and interface inference |
BAMBU_EXTRA_FLAGS |
(empty) | Per-accelerator options, usually set in hw/hls/bambu_args.mk |
BAMBU_TOP |
<name>_bambu_core |
Top function name |
BAMBU_TB |
../tb/<name>_bambu_tb.cpp |
Testbench used by -exe and -sim |
BAMBU_SIMULATOR |
VERILATOR |
Simulator used by -sim |
BAMBU_CXX |
g++ |
Host compiler used by -exe |
The VC707’s xc7vx485t has no Bambu characterization, so the virtex7
default uses the characterized xc7vx690t from the same family. For
UltraScale (virtexu) and UltraScale+ (virtexup) boards, -hls stops with
an error until you set BAMBU_DEVICE to a device that Bambu supports.
Current limitations
- Elements are 32 bits wide and transferred over a 64-bit DMA interface.
- Chunking and batching are not implemented.
- Only the Virtex-7 technology has a default Bambu device mapping.
- Like ESP’s other HLS flows, the Bambu flow is not supported by the Intel/Altera development port.