A collection of CUDA C++ template abstractions for implementing high-performance matrix multiplications and convolutions on NVIDIA GPUs.
NVIDIA CUTLASS — CUDA Templates for High-Performance Linear Algebra
A collection of CUDA C++ template abstractions for implementing high-performance matrix multiplications and convolutions on NVIDIA GPUs.
Agent 可直接安装
这个资产可安装;Agent 先选择当前运行时、检查安装计划,再运行匹配命令。
npx -y tokrepo@latest install 7d20c843-5cea-11f1-9bc6-00163e2b0d79 --target codex先 dry-run 确认安装计划,再运行此命令。
讨论
相关资产
TensorRT — High-Performance Deep Learning Inference by NVIDIA
NVIDIA's SDK for optimizing trained deep learning models for production inference, delivering low latency and high throughput on NVIDIA GPUs through graph optimization, kernel fusion, and precision calibration.
SGLang — Fast LLM Serving with RadixAttention
SGLang is a high-performance serving framework for LLMs and multimodal models. 25.3K+ GitHub stars. RadixAttention prefix caching, speculative decoding, structured outputs. NVIDIA/AMD/Intel/TPU. Apach
cuDF — GPU-Accelerated DataFrame Library by NVIDIA RAPIDS
cuDF is a GPU-accelerated DataFrame library from the NVIDIA RAPIDS suite that provides a pandas-like API for data manipulation at 10-100x the speed on NVIDIA GPUs.
NVIDIA Triton Inference Server — Multi-Framework Model Serving at Scale
Triton Inference Server is NVIDIA's production model serving platform. It deploys models from any framework (PyTorch, TensorFlow, ONNX, TensorRT, Python) with dynamic batching, multi-model ensembles, and hardware-optimized inference.