Skip to content

chore(release): bump t4a CubeCL packages to 0.10.1 - #20

Merged
shinaoka merged 1 commit into
mainfrom
release/v0.10.1
Sep 20, 2026
Merged

shinaoka merged 1 commit into
mainfrom
release/v0.10.1

Conversation

@shinaoka

Copy link
Copy Markdown
Member

Why

The published t4a-cubecl* 0.10.0 crates predate fix(cpp): lower CUDA complex casts with cuComplex helpers (#17). Assign::format_scalar in the published t4a-cubecl-cpp 0.10.0 still emits a plain C++ constructor call for every scalar cast:

match elem {
    Elem::TF32 => write!(f, "nvcuda::wmma::__float_to_tf32({input})"),
    elem => write!(f, "{elem}({input})"),
}

so any kernel that casts a real value into a complex dtype fails NVRTC compilation with

error: no suitable constructor exists to convert from "uint32" to "double2"
  const cuDoubleComplex l_16 = cuDoubleComplex(uint32(0));

This is what tensor4all/tenferro-rs#1833 reports: complex triu/tril, and therefore complex QR, are unavailable on CUDA for every crates.io consumer of tenferro-rs 0.5.0. Repository builds do not see it because tenferro-rs pins this fork by git rev (a2adda17), which already contains the fix.

What

Bump [workspace.package].version and every first-party inter-crate requirement from 0.10.0 to 0.10.1 so the fix can be published. Third-party spin and rand requirements that also read 0.10.0 are left untouched.

After this merges, tagging v0.10.1 publishes the crate set listed in AGENTS.md through .github/workflows/publish.yml, and tenferro-rs can move its pin to =0.10.1.

Verification

  • cargo metadata --no-deps resolves all 13 first-party packages at 0.10.1
  • no source changes; the CUDA complex-cast lowering already on main is what this release carries

🤖 Generated with Claude Code

The published `t4a-cubecl* 0.10.0` crates predate the CUDA complex-cast
lowering fix (`fix(cpp): lower CUDA complex casts with cuComplex helpers`,
#17), so every crates.io consumer of tenferro-rs still hits
`no suitable constructor exists to convert from "uint32" to "double2"`
whenever a kernel casts a real value into a complex dtype
(tensor4all/tenferro-rs#1833).

Bump the workspace package version and every first-party inter-crate
requirement to 0.10.1 so the fix can be published. Third-party `spin` and
`rand` requirements that also read 0.10.0 are left untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant