Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.rst
Original file line number Diff line number Diff line change
Expand Up @@ -133,6 +133,7 @@ Changelog
tensors; use resident export for those models.
- Fix shared ONNX export metadata and Diffusers attention policy: every ``NVFP4QuantExporter`` post-process now upgrades the default-domain opset to at least 23, all FP8 custom-op exports re-run ONNX shape/type inference after setting output metadata, and quantized SDPA derives FP8 MHA enablement from the live Q/K/V quantizers instead of honoring a caller-set ``_disable_fp8_mha`` attribute.
- Fix ONNX FP16 conversion failing to preserve public output types when type inference changes a graph output declaration before output casts are inserted.
- Fix ONNX remote AutoQDQ safety benchmarks returning infinite latency by running the generated engine with ``trtexec_safe`` on the configured target.
- Fix ``examples/hf_ptq/hf_ptq.py`` discarding a completed PTQ run (no checkpoint exported) when the optional post-quantization sanity-check ``generate()`` call raised, for example because ``device_map="auto"`` placed part of the model on CPU. That failure is now caught and only skips the sanity check; export proceeds regardless.
- Fix ``examples/megatron_bridge/export_quantized_megatron_to_hf.py`` storing the MoE router at Megatron's ``moe_router_dtype``, which is a routing *compute* dtype, not a storage one. The router now exports at the export ``dtype`` like every other unquantized weight, matching what ``hf_ptq.py`` and the released NVFP4 checkpoints contain; pass ``moe_router_dtype`` to ``export_mcore_gpt_to_hf`` explicitly if you want the old fp32 storage.
- Fix unified Megatron export writing a second, unreferenced copy of the vocab embedding when a model with MTP layers is exported with pipeline parallelism. The duplicate was never loaded but inflated the checkpoint by the size of the embedding (about 1 GB for Qwen3.6-35B-A3B); re-export to reclaim the space.
Expand Down
10 changes: 8 additions & 2 deletions docs/source/guides/9_autotune.rst
Original file line number Diff line number Diff line change
Expand Up @@ -251,11 +251,17 @@ To use remote autotuning during Q/DQ placement optimization, run with ``trtexec`
**Requirements:**

* TensorRT 10.15 or later
* Valid remote autotuning configuration
* Valid ``ssh://`` remote autotuning configuration without a password
* Non-interactive SSH key authentication from the host to the target
* ``trtexec_safe`` on the target, alongside the configured ``remote_exec_path``
* ``--use_trtexec`` must be set (benchmarking uses ``trtexec`` instead of the TensorRT Python API)
* ``--safe --skipInference`` must be enabled via ``--trtexec_benchmark_args``

Replace ``<remote autotuning config>`` with an actual remote autotuning configuration string (see ``trtexec --help`` for more details). Other TensorRT benchmark options (e.g. ``--timing_cache``, ``--warmup_runs``, ``--timing_runs``, ``--plugin_libraries``) are also available; run ``--help`` for details.
Replace ``<remote autotuning config>`` with an actual remote autotuning configuration string (see ``trtexec --help`` for more details). ModelOpt uses the configuration to build the engine with remote autotuning, copies the generated engine to an internal temporary path on the target, and runs ``trtexec_safe`` there to measure GPU compute time. ModelOpt attempts to remove the temporary engine after each benchmark.

Other TensorRT benchmark options (e.g. ``--timing_cache``, ``--warmup_runs``, ``--timing_runs``, ``--plugin_libraries``) are also available; run ``--help`` for details.

``--plugin_libraries`` applies to the host-side engine build only. ModelOpt does not transfer or load custom plugin libraries on the remote target, so this workflow does not support plugin-dependent remote safety engines.

Low-Level API Usage
===================
Expand Down
11 changes: 8 additions & 3 deletions examples/onnx_ptq/autotune/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -247,12 +247,17 @@ python3 -m modelopt.onnx.quantization.autotune \
**Requirements:**

- TensorRT 10.15 or later
- Valid remote autotuning configuration
- Valid `ssh://` remote autotuning configuration without a password
- Non-interactive SSH key authentication from the host to the target
- `trtexec_safe` on the target, alongside the configured `remote_exec_path`
- `--use_trtexec` must be set (benchmarking uses `trtexec` instead of the TensorRT Python API)
- `--safe --skipInference` must be enabled via `--trtexec_benchmark_args`

Replace `<remote autotuning config>` with an actual remote autotuning configuration string (see `trtexec --help` for more details).
Other TensorRT benchmark options (e.g. `--timing_cache`, `--warmup_runs`, `--timing_runs`, `--plugin_libraries`) are also available; run `--help` for details.
Replace `<remote autotuning config>` with an actual remote autotuning configuration string (see `trtexec --help` for more details). ModelOpt uses the configuration to build the engine with remote autotuning, copies the generated engine to an internal temporary path on the target, and runs `trtexec_safe` there to measure GPU compute time. ModelOpt attempts to remove the temporary engine after each benchmark.

Other TensorRT benchmark options (e.g. `--timing_cache`, `--warmup_runs`, `--timing_runs`, `--plugin_libraries`) are also available; run `--help` for details.

`--plugin_libraries` applies to the host-side engine build only. ModelOpt does not transfer or load custom plugin libraries on the remote target, so this workflow does not support plugin-dependent remote safety engines.

## Programmatic API Usage

Expand Down
Loading
Loading