Developing ASICs is a complex and computationally demanding process that involves multiple design stages and chip design tools needed to transform an RTL description into a physical chip layout. The OpenROAD Initiative aims to improve the automated digital design tooling and make it accessible to a broader community of ASIC engineers, fostering innovation in the domain and opening opportunities for optimizing the digital design process.
Global placement is usually a quite time-consuming step in the RTL-GDSII flow as it’s responsible for solving complex, non-linear optimization problems, being specifically demanding for very-large-scale integration (VLSI) designs. At the same time, global placement is highly parallelizable, bearing a huge potential for GPU acceleration.
Antmicro, a Principal Member of the OpenRoad Initiative, having years of experience in enhancing and customizing open source ASIC development tooling and developing custom ASIC and FPGA IP cores for its customers, has been spearheading a significant advancement within this global placement stage - the implementation of an improved GPU acceleration process in OpenROAD’s global placement module (GPL).
The core improvement lies in accelerating the GPL through parallel processing, which makes it possible to execute workload-heavy tasks on GPU or run them concurrently on multiple CPU cores. This implementation also brings more hardware flexibility, achieved via migration from CUDA to Kokkos, which allows using both CPU- and GPU-based parallel execution on various parallel computation backends without modifying the underlying code. In this blog, we describe the implemented Kokkos-based GPU acceleration of global placement and illustrate the efficiency gains it brings to the RTL-GDSII design flow.
Improved GPU acceleration with Kokkos
The global placement module in OpenROAD was previously rewritten to enable GPU acceleration using NVIDIA’s CUDA Toolkit. As the implemented GPU variant was NVIDIA-specific, switching to another GPU provider presupposed rewriting the GPL code to use other vendor-specific APIs.
Kokkos, a C++ programming system for writing performance-optimized applications, helps to address this challenge of vendor dependency. It offers a unified API interface for both computation as well as data allocation and layout, enabling developers to write a single code that can be later compiled for a specific backend. This results in massive performance gains and is much more efficient from the perspective of code maintenance and reusability.
How the Kokkos implementation works
The Kokkos API includes the machine model that has two crucial components: a memory space, which defines how memory is allocated, and the execution space, which specifies how to parallelize work using one or more memory spaces. This model enables portability across various devices that have different memory layouts or need to differently parallelize work for the highest performance. In addition to the machine model, Kokkos also provides the programming model that allows for compile-time transformation of algorithms for a given backend. In combination, these abstractions allow mapping generic algorithms into device-specific technology.
Ensuring determinism in global placement
The GPL makes extensive use of floating-point numbers, which can produce different results when performed in a different order, resulting in non-determinism. For example, specific operations in the GPL require summing a large array. A slightly different result can then be used in subsequent calculations, with these small differences potentially accumulating into a larger error that could meaningfully affect the output.
Initially, the new implementation of GPU acceleration in the GPL resulted in unprecedentedly fast performance. However, as OpenROAD tooling is designed to deliver stable, deterministic results whenever possible, some performance trade-offs were required in order to preserve this behavior. In particular, certain operations had to be performed in a single thread, rather than concurrently. Even so, the final implementation still delivers impressive efficiency gains, as demonstrated in the following section.
Benchmarking the performance of GPU acceleration in global placement with the new Kokkos-based CUDA backend
We have tested the new Kokkos-based implementation across different backends on NVIDIA’s GeForce RTX 3090 GPU and Intel Core i9-12900 CPU, using the ariane133 design:
These results demonstrate that the newly implemented Kokkos-based CUDA backend brings substantial runtime improvements compared to the original placer without GPU parallelization.
Moreover, when compared with the existing CUDA implementation for GPU acceleration, the Kokkos-based CUDA backend maintains near-native performance while providing much-needed flexibility and vendor-agnosticism:
In the context of GPL improvements in OpenROAD, there’s another important update that’s currently in the works - migration from OpenMP to Kokkos, as well as enabling Kokkos by default in the CMake and Bazel builds of OpenROAD. Once finished, this work will mark the completion of Kokkos enablement in OpenROAD - keep an eye on the upcoming blog notes where we’ll share more details about the implementation and its outcomes.
Accelerate your digital design development flow with Antmicro
Implemented with the crucial contribution of Antmicro, the GPU acceleration in OpenROAD’s global placement module brings enhanced performance to the RTL-to-GDS flow, also enabling portability across different GPU vendors. Thanks to the underlying Kokkos framework, this functionality can be leveraged both on CPU and GPU backends, ultimately reducing execution time for OpenROAD-based ASIC development flows.
Antmicro has years of experience in enhancing and customizing open source ASIC development tooling for its customers, with other notable examples including the development of a new resynthesis strategy based on simulated annealing, implementation of automatic clock gating, and the performance optimizations in the OpenROAD-flow-scripts.
Similarly, we can help you customize and adopt RTL-to-GDS tooling to your specific requirements and use cases. If you want to benefit from our expertise in the ASIC/FPGA domain to build and customize digital designs for your products, automate the generation of top-level designs with our design aggregation toolkit Topwrap, and test how your ASIC chips would work in real-life conditions by integrating Verilator into your processes, reach out to us at contact@antmicro.com.