Skip to content

Design a C/C++ Accelerator with PandA Bambu HLS

Last updated

This guide shows how to design an ESP accelerator in C/C++ and synthesize it with PandA Bambu, an open-source high-level synthesis (HLS) tool. The flow follows the same structure as ESP’s other accelerator flows: generate a skeleton, implement and verify the kernel, install the generated RTL, and integrate the accelerator into an SoC.

The Bambu HLS flow is available on the dev branch and is planned for the next ESP release. Commands and options may change before the release. Coming soon

Platform support: the flow targets AMD/Xilinx FPGA technologies. For the Virtex-7 technology library (virtex7, used by the VC707 and proFPGA XC7V2000T), ESP selects a Bambu device characterization automatically. For other technologies, set BAMBU_DEVICE as described in Configuration variables.

This guide assumes that you are familiar with the ESP infrastructure and know how to configure, simulate and prototype a single-core SoC.

Prerequisites

  • An ESP checkout on the dev branch, set up as described in the setup guide.
  • PandA Bambu built from the feature/unified_ac_channels branch. ESP validation used revision 4150ff6e20d543eb5d241df7085046b6a355f283, Clang 16 and Verilator 4.038.
  • For Bambu RTL co-simulation, the selected simulator (Verilator by default) and 32-bit C++ linking support, such as the g++-multilib package on Ubuntu.
  • (optional) Icarus Verilog (iverilog and vvp) for the standalone DMA test of the example accelerator.

Point ESP at the Bambu installation. BAMBU_ENV is optional and names a settings script that ESP sources before each Bambu invocation:

export BAMBU=<bambu_path>/bin/bambu
export BAMBU_ENV=<bambu_path>/settings.sh

If bambu is already on your PATH and needs no settings script, you can omit both variables.

Generate the accelerator skeleton

Run ESP’s interactive accelerator generator from the ESP root and select B for Bambu HLS:

cd <esp>
./tools/accgen/accgen.sh
=== Initializing ESP accelerator template ===

  * Enter accelerator name [dummy]: example
  * Select design flow (Stratus HLS, Vivado HLS, hls4ml, Catapult HLS, RTL, Bambu HLS) [S]: B
  * Enter ESP path [<esp>]:
  * Enter unique accelerator id as three hex digits [04A]: 07B
  * Enter accelerator registers
    - register 0 name [size]: len
    - register 0 default value [1]: 64
    - register 0 max value [64]: 1024
    - register 1 name []:
  ...

The Bambu flow currently supports a subset of the generator options. The generator warns and applies these settings automatically:

Option Bambu flow value
Data bit-width 32 bits (32-bit elements over a 64-bit DMA interface)
Chunking factor 1 (no chunking)
Batching factor 1 (no batching)

The accelerator is created in accelerators/bambu_hls/<name>_bambu/:

<name>_bambu/
├── hw
│   ├── <name>_bambu.xml           # accelerator description and registers
│   ├── hls
│   │   ├── Makefile -> ../../../common/hls/Makefile
│   │   └── bambu_args.mk          # per-accelerator Bambu flags (optional)
│   ├── src
│   │   ├── <name>_bambu.cpp       # C/C++ kernel
│   │   └── <name>_bambu_basic_dma64/
│   │       └── <name>_bambu_basic_dma64.v   # ESP socket wrapper
│   └── tb
│       └── <name>_bambu_tb.cpp    # C++ testbench
└── sw
    ├── baremetal                   # bare-metal test application
    └── linux                       # device driver and user-space application

Implement the accelerator

The generated kernel in hw/src/<name>_bambu.cpp is a dataflow design with concurrent load, compute and store stages connected by ping-pong buffers. The placeholder computation squares each input element; replace it with your accelerator’s function. Each configuration register becomes a scalar argument of the top function and a conf_info_<register> port of the wrapper.

Keep these conventions:

  • The C/C++ top function must be named <name>_bambu_core. Bambu generates a Verilog module with that name, and the ESP socket wrapper instantiates it.
  • The wrapper in hw/src/<name>_bambu_basic_dma64/ adapts the generated core’s AXI-Stream channels to the ESP accelerator socket. Only a 64-bit DMA wrapper is provided.
  • Register handling, interrupts, scatter-gather address translation and DMA over the network on chip are provided by the generated ESP socket, not by the Bambu core.
  • Update the testbench in hw/tb/<name>_bambu_tb.cpp so that it checks your outputs and prints PASS on success; the co-simulation target looks for that string.

Add extra Bambu options for one accelerator in hw/hls/bambu_args.mk, for example:

BAMBU_EXTRA_FLAGS = --no-clean

Synthesize and verify

Run the Bambu targets from the working directory of your SoC. The targets are named after the accelerator directory:

cd <esp>/socs/xilinx-vc707-xc7vx485t

# Native C++ execution of the testbench (no HLS)
make <name>_bambu-exe

# High-level synthesis and installation into ESP
make <name>_bambu-hls

# Bambu RTL co-simulation of the testbench
make <name>_bambu-sim
  • -exe compiles the testbench and kernel with the host C++ compiler, using Bambu’s headers.
  • -hls runs Bambu in hw/hls-work-<tech>/ and installs the accelerator XML, the socket wrapper, the generated core and its memory initialization files in tech/<tech>/acc/<name>_bambu/. The log is saved in the SoC working directory’s HLS log folder.
  • -sim re-runs Bambu with its generated testbench and simulator, and passes only if the testbench log contains PASS.

Other targets:

Target Description
<name>_bambu-clean Remove Bambu outputs from the work directory
<name>_bambu-distclean Also remove the installed RTL from tech/
bambuhls_acc Run HLS for every Bambu accelerator
bambuhls_acc-clean, bambuhls_acc-distclean Clean every Bambu accelerator
print-available-acc List accelerators, including Bambu accelerators

Bambu accelerators are not built by the Stratus HLS make acc target; use bambuhls_acc instead.

Integrate the accelerator into an SoC

After make <name>_bambu-hls, the accelerator appears in the ESP configuration GUI like any other installed accelerator:

  1. Run make esp-xconfig and place the accelerator in a tile with the basic_dma64 implementation. Use a configuration with 64-bit DMA, such as the default Ariane configuration.
  2. Generate the sockets and build the bare-metal test:

    make socketgen
    make <name>_bambu-baremetal
    
  3. Run a full-system RTL simulation with the bare-metal test:

    make sim TEST_PROGRAM=$PWD/soft-build/ariane/baremetal/<name>_bambu.exe
    
  4. Build the bitstream and test on FPGA with the usual single-core SoC targets:

    make vivado-syn
    make fpga-program
    make fpga-run
    

ModelSim and Vivado compile Bambu accelerator libraries automatically and stage their memory initialization files. When simulating with ModelSim, use the same version that compiled the cached AMD/Xilinx simulation libraries. For an existing Vivado project, run make vivado-update to pick up a new accelerator.

Example: sqr_bambu

accelerators/bambu_hls/sqr_bambu/ is a validated example. It squares unsigned 32-bit words modulo 232 with load, compute and store stages that run concurrently, using two 16-word banks at each stage boundary.

Register Offset Description
len 0x40 Number of 32-bit words; a positive multiple of 16
base_in 0x44 Input offset, in 64-bit DMA beats
base_out 0x48 Output offset, in 64-bit DMA beats

Use disjoint input and output regions. The hardware does not reject invalid lengths, so software must enforce them.

The reference SoC has one Ariane processor tile, one memory tile, one I/O tile and one SQR_BAMBU accelerator tile with the basic_dma64 implementation, caches disabled and 64-bit DMA. Instead of using the GUI, you can add the tile to socgen/esp/.esp_config directly:

TILE_1_0 = 2 acc SQR_BAMBU basic_dma64 0 0 sld

Build and validate it from the ESP root:

make -C socs/xilinx-vc707-xc7vx485t sqr_bambu-hls
make -C socs/xilinx-vc707-xc7vx485t sqr_bambu-sim

# Standalone DMA test of the installed RTL (requires Icarus Verilog)
bash accelerators/bambu_hls/sqr_bambu/hw/tb/run_dma_test.sh

make -C socs/xilinx-vc707-xc7vx485t esp-xconfig
make -C socs/xilinx-vc707-xc7vx485t socketgen sqr_bambu-baremetal
make -C socs/xilinx-vc707-xc7vx485t sim \
  TEST_PROGRAM="$PWD/socs/xilinx-vc707-xc7vx485t/soft-build/ariane/baremetal/sqr_bambu.exe"
make -C socs/xilinx-vc707-xc7vx485t vivado-syn

The bare-metal application runs three 272-word invocations and checks the results. A passing run prints:

sqr_bambu PASS total_errors=0

ESP’s simulation harness then prints Failure: Program Completed! as its normal end-of-test assertion. Check the application’s PASS line rather than that assertion.

Verification note: sqr_bambu-sim sets BAMBU_SKIP_VERIFICATION because Bambu’s sequential C model cannot represent the concurrent bank handoff. The testbench still checks the RTL outputs independently, and the standalone DMA test covers descriptors, stalls and repeated starts.

The example includes a bare-metal application only; it has no Linux driver or user-space application.

Configuration variables

Set these variables in the environment or on the make command line to override the defaults:

Variable Default Description
BAMBU bambu Bambu executable, on PATH or as a full path
BAMBU_ENV (empty) Settings script sourced before running Bambu
BAMBU_DEVICE xc7vx690t,-3,ffg1930 for virtex7 Bambu device as device,speedgrade,package; required for other technologies
BAMBU_CLOCK_PERIOD 10 Target clock period in ns
BAMBU_FLAGS see accelerators/bambu_hls/common/hls/Makefile Base Bambu options, including Clang 16 and interface inference
BAMBU_EXTRA_FLAGS (empty) Per-accelerator options, usually set in hw/hls/bambu_args.mk
BAMBU_TOP <name>_bambu_core Top function name
BAMBU_TB ../tb/<name>_bambu_tb.cpp Testbench used by -exe and -sim
BAMBU_SIMULATOR VERILATOR Simulator used by -sim
BAMBU_CXX g++ Host compiler used by -exe

The VC707’s xc7vx485t has no Bambu characterization, so the virtex7 default uses the characterized xc7vx690t from the same family. For UltraScale (virtexu) and UltraScale+ (virtexup) boards, -hls stops with an error until you set BAMBU_DEVICE to a device that Bambu supports.

Current limitations

  • Elements are 32 bits wide and transferred over a 64-bit DMA interface.
  • Chunking and batching are not implemented.
  • Only the Virtex-7 technology has a default Bambu device mapping.
  • Like ESP’s other HLS flows, the Bambu flow is not supported by the Intel/Altera development port.