baracuda-flashinfer
v0.0.1-alpha.77 ExperimentalSafe, typed Rust wrappers for NVIDIA FlashInfer's inference-serving kernels: batched paged-KV attention decode, decode-time KV-cache append, cascade / prefix-cache attention-state merge, and sort-free top-K / top-P / min-P sampling. The canonical vLLM-style serving surface for the baracuda CUDA stack. Apache-2.0 (FlashInfer upstream).
Quick Verdict
- โActively maintained (updated 4d ago)
- !Pre-1.0: API may have breaking changes
- โPermissive license (MIT OR Apache-2.0)
Security
Deep Insights
188 downloads in the last 30 days (6/day), up 236% from the previous period.
The primary maintainer publishes 161 crates. This suggests deep Rust expertise and long-term commitment to the ecosystem.
At 19KB, baracuda-flashinfer is lightweight. Small crate size correlates with focused, well-scoped functionality.
Health Breakdown
Recency, release consistency, active ratio
Yanked ratio, deps, size, maturity, features
Reverse deps, ownership, ecosystem
Downloads, momentum, growth trend
Docs, repo, license, metadata