ONNX Runtime is a high-performance inference engine for machine learning models in the ONNX format. Developed by Microsoft, it accelerates model serving across CPU, GPU, and specialized hardware with a unified API for Python, C++, C#, Java, and JavaScript.
ONNX Runtime — Cross-Platform ML Model Inference Engine
ONNX Runtime is a high-performance inference engine for machine learning models in the ONNX format. Developed by Microsoft, it accelerates model serving across CPU, GPU, and specialized hardware with a unified API for Python, C++, C#, Java, and JavaScript.
Agent 可直接安装
这个资产可安装;Agent 先选择当前运行时、检查安装计划,再运行匹配命令。
npx -y tokrepo@latest install 0e90de1c-3d9d-11f1-9bc6-00163e2b0d79 --target codex先 dry-run 确认安装计划,再运行此命令。
讨论
相关资产
NVIDIA Triton Inference Server — Multi-Framework Model Serving at Scale
Triton Inference Server is NVIDIA's production model serving platform. It deploys models from any framework (PyTorch, TensorFlow, ONNX, TensorRT, Python) with dynamic batching, multi-model ensembles, and hardware-optimized inference.
ONNX Runtime — Cross-Platform ML Inference Accelerator
ONNX Runtime is Microsoft's high-performance inference engine for machine learning models in the ONNX format. It supports CPU, GPU, and specialized hardware accelerators across Linux, Windows, macOS, iOS, Android, and the web browser.
ONNX Runtime — Cross-Platform ML Inference Accelerator
A high-performance inference engine for ONNX models that runs on CPU, GPU, and specialized hardware across cloud, edge, and mobile.
ONNX Runtime — Cross-Platform ML Inference and Training Accelerator
High-performance inference engine for ONNX models across CPUs, GPUs, and edge devices with broad framework support.