Back to Blog

Integrating Tachyon's Cryptography with Zakura Common

We're beginning to upstream Tachyon's recursive proofs into the core of Zcash's newest cryptography stack.
An illustrated black bowl of udon noodles patterned with pink sakura blossoms.
Udon is a cryptography backend being integrated into Zakura Common.

We recently joined forces with Valar Group to develop Zakura, a new full node for Zcash. As part of that work we built Common, a new cryptography stack for Zcash that now powers Zakura and many other projects in the Zcash ecosystem. Common brought substantial performance improvements, including dramatically faster proving for the Ironwood zk-SNARK.

Tachyon's recursive proofs would ordinarily be more expensive to build, demand more memory, and produce larger binaries. Because the protocol deliberately reuses much of Ironwood's cryptography, it addresses those costs with Common's optimized codebase. But Tachyon also extends that cryptography, and some of the bottlenecks it exposes limit Ironwood proving too.

Today we are beginning the process of upstreaming Tachyon's recursive proofs into the core of Zakura Common. This requires an architectural adjustment to parts of the stack. As a first step we are releasing several new Zakura crates that subsume mature components of our Ragu codebase, removing the shims and compatibility layers that previously kept those components from benefiting from Common's optimized kernels.

Shared Foundations

Tachyon and Ironwood both use the Pasta elliptic curve cycle for ECDLP-based zk-SNARKs. Zcash designed Ironwood (Orchard) around these curves deliberately, anticipating that a future upgrade like Tachyon would need them. But Tachyon's proving requirements are not identical to Ironwood's:

  • Its vector commitments need longer generator vectors that partially overlap with Ironwood's.
  • Recursion means we prove over both curves in the cycle, not just one.
  • Both use Poseidon, but Tachyon uses a slightly different parameterization.

Bringing these requirements into Common's arithmetic interfaces lets us share applicable optimizations and avoid unnecessary code duplication.

Udon 🍜

Today, we're releasing Udon (zakura-udon), a new Rust crate that replaces Common's pasta_curves fork and much of the arithmetic in halo2_proofs, the Halo 2 zk-SNARK prover used by Ironwood. Udon is a low-level implementation of the Pasta curves and fields which inherits most of the optimizations in Common, while leaving behind a few that were not (yet) worth their complexity.

Resolving Technical Debt

Halo 2 was designed with the Pasta curves in mind, and it makes a number of assumptions that only hold for them in practice. Nevertheless, its code is written against the generic ff/group traits from zkcrypto, which have no room for curve-specific detail.

The problem begins with pasta_curves, which centers its entire API around these generic traits. Optimized, low-level APIs are not accessible through those interfaces since they are not defined by pasta_curves in the first place. In some cases, pasta_curves must compensate for this with its own extension traits.

Generic APIs are still necessary, at some layer, because both curves have a vast number of generic algorithms defined equally for them. But many of those algorithms are also heavily parallelized, and pasta_curves did not want to presume a particular execution model. This left halo2_proofs and other crate consumers to define generic algorithms that they could not possibly optimize as well as pasta_curves itself.

Once Common began adding aggressive Pasta-specific optimizations, every substantive performance change had to tunnel through one of three mechanisms:

  • Extension traits turned into backend traits. CurveExt grew methods for multiexponentiation, point FFTs, shared-scalar batch MSMs and prepared zero checks, with generic fallbacks so that the trait stays nominally implementable by any curve.
  • Bridge functions. pasta_curves exports #[doc(hidden)] monomorphic functions (partial Montgomery reductions, dedicated squaring chains, an unchecked two-lane mixed addition) as internal cross-crate bridges for halo2_proofs. They are consumed from halo2_gadgets, orchard and sinsemilla as well.
  • Reflection. Many hot paths recover the concrete types at runtime through TypeId and dyn Any downcasts so they can call Pasta-specific routines the generic traits cannot express, and halo2_proofs contains closed-world enums with Pallas and Vesta variants inside modules that are generic over F: Field. There are many such dispatch sites in halo2_proofs, each keeping a generic fallback that never runs, and the same pattern is repeated in halo2_gadgets.

Udon integrates these layers together. It owns all of the algebra currently split between pasta_curves and halo2_proofs: the field and curve traits, their arithmetic, and the algorithms built on them. It is openly specialized to Pallas and Vesta. Because one crate sees every representation detail, the escape hatches become ordinary monomorphic calls and the duplicated kernels collapse to one implementation each.

Execution and Memory

Udon decides how to do arithmetic efficiently within the application's constraints. The application provides those constraints, as well as the runtime and working storage. Udon is independent of the specific backend (such as Rayon) used for parallelism, and the caller owns the buffers.

This split is deliberate: the caller is the one that knows which operations are independent, which results it needs next, and which buffers it can afford to keep. It expresses its resource choices through the buffers, limits, and executor it passes to an operation.

As an example, a forward FFT over a caller-owned buffer looks like this:

let transform = Transform::new(Domain::new(k).unwrap().subgroup());
transform.forward(&mut coeffs, options, &executor, &mut scratch).unwrap();

The buffer, the executor, and the scratch space all belong to the caller. Udon's runtime arithmetic is no_std and does not allocate: it never grows a working buffer or creates a thread on its own. The application can allocate during setup and reuse the same workspace across operations and proving rounds.

This is done through "plans." An FftPlan or MsmPlan resolves an arithmetic strategy, reports its storage requirements, and can be retained for later inputs while the working buffers change. Direct calls such as forward select a strategy using the scratch supplied for that call; retaining a plan lets the application reuse those planning decisions.

Executors and Limits

The Executor interface supplies a fork-join operation. Udon ships a serial implementation, and an adapter can connect it to Rayon's work-stealing pool. The application chooses and owns the pool.

ExecutionOptions supplies the concurrency and workspace constraints used to plan an operation:

let options = ExecutionOptions::default()
    .with_task_budget(TaskBudget::new(8).unwrap())
    .with_memory_limit(64 * 1024 * 1024);

The default budget is one task. Larger budgets permit more parallel work, while input size, algorithm structure, and available workspace decide how much is useful. Planning fails if no strategy fits the supplied limits.

Incremental Execution

Scoped calls return after an entire operation finishes. For provers that need finer control, Udon exposes FFTs and MSMs as incremental runs. A polynomial expansion, for example, produces blocks of evaluations on the cosets that make up a larger evaluation domain. One completed block, called a residue, can be consumed while the other blocks are still computing.

Interpolation can similarly release an input once its last contribution has been absorbed, and an MSM can begin on scalar fragments as soon as their values are final. The application can schedule these tasks alongside its own computation, controlling priorities and the number of tasks in flight.

This gives a prover control below the level of a whole FFT or MSM. It can favor work that releases a large intermediate, overlap independent stages, and reuse storage across proving rounds without reimplementing the arithmetic.

Artifacts

Udon has a compile-time companion crate, Bento (zakura-bento). Bento defines a plain-data interface for embedding prepared data in a binary with layout, alignment, and encoding accounted for automatically.

Udon's field elements and affine points implement that interface in their native representation, so an embedded fixed-base table or many other kinds of rich data structures can be borrowed directly at runtime: no decoding pass, and no redundant allocations.

Common already embeds artifacts in its binaries to reduce warm-up cost, but still pays for redundant allocations when it loads them. Bento keeps the embedding and removes that second cost.

Variable-time Arithmetic

Neither halo2_proofs nor pasta_curves ever offered strong side-channel resistance; the cost across an entire proof is too high. Yet both inherited from ff and group a philosophy for API design that required all behavior to be maximally side-channel resistant unless methods carried an explicit vartime label.

This noble default was the wrong approach for a proving stack. Provers and circuits need to be fast and easy to write, and constant-time arithmetic was never realistically available in that setting; provers paid for constant-time arithmetic even as variable-time operations were performed anyway. Further, numerous arbitrary operations (especially operator overloading) paid a performance penalty due to the inability to ergonomically express intentions.

Udon inverts the default. All of its field and curve arithmetic is variable-time, and it makes no guarantee about timing or data-dependent memory access. Instead, code is explicitly annotated as side-channel resistant and those particular areas (or crates) are expected to uphold that contract and maintain it as the codebase(s) change.