fix: remove tie weight check by RayenTian · Pull Request #700 · NVIDIA-NeMo/RL

RayenTian · 2025-07-21T05:40:02Z

What does this PR do ?

This pull request removes the tie weight check in the DTensor Worker, addressing a previous limitation with training models that utilize tie_word_embeddings when tensor parallel size (tp_size) > 1.

Issues

closes #684

Test Result

For remove `NRL_SKIP_TIED_WEIGHT_CHECK` env variable

Experiments were run with meta-llama/Llama-3.2-1B, Qwen/Qwen2.5-1.5B-Instruct and google/gemma-2-2b-it. With the truncate rate set within a reasonable range, both tp_size=1 and tp_size=2 produced comparable reward results, indicating that training proceeds as expected. This demonstrates consistency in model behavior across different tensor parallel configurations, confirming correct handling of tied word embeddings during training.

Main setup:

policy.generation.temperature=1.0
policy.max_total_sequence_length=2048
cluster.gpus_per_node=8

Llama-3.2-1B

Qwen2.5-1.5B-Instruct

google/gemma-2-2b-it

Test with:

policy.dynamic_batching.enabled=true
policy.sequence_packing.enabled=false
policy.train_global_batch_size=128

For remove `model.config.tie_word_embeddings` in `nemo_rl/models/dtensor/parallelize.py`

Experiments were conducted using both the Qwen/Qwen2.5-1.5B-Instruct and Qwen/Qwen2.5-7B-Instruct models, representing the original tied and untied weight configurations, respectively. I disabled Qwen's optimized parallel plan in PARALLIZE_FUNCTIONS, forcing both models to utilize the HP plan for testing.

The results indicate that, after removing the model.config.tie_word_embeddings check, the HP plan and the optimized plan produce identical outcomes.

Qwen/Qwen2.5-1.5B-Instruct

This model uses tied word embeddings by default.
Current results are from a short test of 10 steps.

Qwen/Qwen2.5-7B-Instruct

This model uses untied word embeddings by default.
Current results are from a short test of 10 steps.

More Custom Test

Some additional tests were also conducted here.
For models like Qwen/Qwen2.5-1.5B-Instruct, which by default use tied word embeddings, we forcibly disabled the tie_word_embedding setting. The results showed significant anomalies in both the token_mult_prob_error and the reward. We suspect this may be because vLLM still keeps the embeddings tied on its side, so when you call update_weights, there might be some issues (in practice, vLLM may be using either the embed or the lm_head weight, rather than truly untying them as in training). This causes the token_mult_prob_error to behave abnormally.

For models like Qwen/Qwen2.5-7B-Instruct, which by default do not use tied embeddings, we forcibly enabled tie_word_embedding. The results showed large differences in reward, but the token_mult_prob_error remained stable. This might be because, originally, the embed and lm_head weights were different due to being untied; forcing a tie essentially removes the original lm_head weights. When we call update_weights to vLLM, its lm_head is also overwritten with the embed weights, so the token_mult_prob_error remains normal. However, since we forcibly replaced the original weights, the reward becomes abnormal.

Additional Notes on Gemma Model Testing

gemma-3-1b-it: This model sets kv_head=1, which is currently not compatible with tensor parallelism (TP=2). As a result, it could not be tested in a TP=2 configuration.

yuki-97 · 2025-07-22T06:03:26Z

Thanks @RayenTian for verifying and removing these things!

Can you also check if we can also remove model.config.tie_word_embeddings in HF TP plan in nemo_rl/models/dtensor/parallelize.py?

yuki-97 · 2025-07-22T07:14:03Z

One more thing, can you also search issues/227 in the repo and remove (or update) them?

SahilJain314 · 2025-07-23T21:17:42Z

can you please attach the token_mult_prob_error plots to the description? (we can't always see the errors when we just look at convergence plots)

RayenTian · 2025-07-24T05:22:01Z

can you please attach the token_mult_prob_error plots to the description? (we can't always see the errors when we just look at convergence plots)

Thank you for the suggestion. I have added the plots as requested.

SahilJain314

thanks for rebasing, lgtm now.

Signed-off-by: ruit <ruit@nvidia.com>

…h are not used anymore Signed-off-by: ruit <ruit@nvidia.com>

Signed-off-by: ruit <ruit@nvidia.com>

RayenTian added the CI:L1 Run doctests, unit tests, and functional tests label Jul 21, 2025

RayenTian temporarily deployed to nemo-ci July 21, 2025 05:40 — with GitHub Actions Inactive

RayenTian added CI:L0 Run doctests and unit tests and removed CI:L1 Run doctests, unit tests, and functional tests labels Jul 21, 2025

RayenTian temporarily deployed to nemo-ci July 21, 2025 08:20 — with GitHub Actions Inactive

RayenTian temporarily deployed to public July 21, 2025 08:21 — with GitHub Actions Inactive

RayenTian requested review from joyang-nv and yuki-97 July 21, 2025 09:17

RayenTian added CI:L0 Run doctests and unit tests and removed CI:L0 Run doctests and unit tests labels Jul 22, 2025

RayenTian temporarily deployed to nemo-ci July 22, 2025 03:16 — with GitHub Actions Inactive

RayenTian force-pushed the ruit/remove_tie_weight_check branch from e610b0e to a4dbcec Compare July 22, 2025 04:32

RayenTian temporarily deployed to public July 22, 2025 04:34 — with GitHub Actions Inactive

RayenTian added CI:L1 Run doctests, unit tests, and functional tests and removed CI:L0 Run doctests and unit tests labels Jul 22, 2025

RayenTian temporarily deployed to nemo-ci July 22, 2025 05:05 — with GitHub Actions Inactive

yuki-97 reviewed Jul 22, 2025

View reviewed changes

Comment thread nemo_rl/models/huggingface/common.py Outdated

Comment thread nemo_rl/models/policy/dtensor_policy_worker.py Outdated

RayenTian force-pushed the ruit/remove_tie_weight_check branch from a4dbcec to b1cd9b6 Compare July 23, 2025 01:50

RayenTian temporarily deployed to public July 23, 2025 01:52 — with GitHub Actions Inactive

RayenTian force-pushed the ruit/remove_tie_weight_check branch from b1cd9b6 to fea99a2 Compare July 24, 2025 03:21

github-actions Bot added the Documentation Improvements or additions to documentation label Jul 24, 2025

RayenTian temporarily deployed to public July 24, 2025 03:23 — with GitHub Actions Inactive

RayenTian temporarily deployed to public July 24, 2025 03:27 — with GitHub Actions Inactive

RayenTian commented Jul 24, 2025

View reviewed changes

Comment thread docs/model-quirks.md

RayenTian force-pushed the ruit/remove_tie_weight_check branch from c6386b5 to d741557 Compare July 29, 2025 03:26

RayenTian temporarily deployed to public July 29, 2025 03:28 — with GitHub Actions Inactive

RayenTian added the CI:L0 Run doctests and unit tests label Aug 5, 2025

RayenTian temporarily deployed to nemo-ci August 5, 2025 09:09 — with GitHub Actions Inactive

RayenTian temporarily deployed to nemo-ci August 5, 2025 10:19 — with GitHub Actions Inactive

RayenTian added CI:L0 Run doctests and unit tests and removed CI:L0 Run doctests and unit tests labels Aug 6, 2025

RayenTian temporarily deployed to nemo-ci August 6, 2025 01:20 — with GitHub Actions Inactive

RayenTian temporarily deployed to nemo-ci August 6, 2025 02:21 — with GitHub Actions Inactive

RayenTian force-pushed the ruit/remove_tie_weight_check branch from c9f41b3 to 292cc83 Compare August 6, 2025 07:15

RayenTian added CI:L0 Run doctests and unit tests and removed CI:L0 Run doctests and unit tests labels Aug 6, 2025

RayenTian temporarily deployed to nemo-ci August 6, 2025 07:16 — with GitHub Actions Inactive

RayenTian temporarily deployed to public August 6, 2025 07:17 — with GitHub Actions Inactive

terrykong previously approved these changes Aug 8, 2025

View reviewed changes

SahilJain314 previously approved these changes Aug 8, 2025

View reviewed changes

RayenTian added 5 commits August 8, 2025 22:39

remove tie weight check

5bcf392

Signed-off-by: ruit <ruit@nvidia.com>

remove SKIP_DTENSOR_TIED_WEIGHTS_CHECK and self.num_tied_weights whic…

95fb484

…h are not used anymore Signed-off-by: ruit <ruit@nvidia.com>

remove model.config.tie_word_embeddings

b860e9c

Signed-off-by: ruit <ruit@nvidia.com>

modify doc for tie weight

ba53b32

Signed-off-by: ruit <ruit@nvidia.com>

docs: remove tied weights section for Gemma-3 models

a5d70d8

Signed-off-by: ruit <ruit@nvidia.com>

terrykong dismissed stale reviews from SahilJain314, yuki-97, and themself via a5d70d8 August 8, 2025 22:40

terrykong force-pushed the ruit/remove_tie_weight_check branch from 292cc83 to a5d70d8 Compare August 8, 2025 22:40

terrykong enabled auto-merge August 8, 2025 22:40

terrykong approved these changes Aug 8, 2025

View reviewed changes

terrykong temporarily deployed to public August 8, 2025 22:42 — with GitHub Actions Inactive

terrykong added this pull request to the merge queue Aug 8, 2025

Merged via the queue into main with commit fecf71e Aug 9, 2025
23 checks passed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

fix: remove tie weight check#700

fix: remove tie weight check#700
terrykong merged 5 commits intomainfrom
ruit/remove_tie_weight_check

RayenTian commented Jul 21, 2025 •

edited

Loading

Uh oh!

Uh oh!

Uh oh!

yuki-97 commented Jul 22, 2025

Uh oh!

yuki-97 commented Jul 22, 2025

Uh oh!

SahilJain314 commented Jul 23, 2025

Uh oh!

Uh oh!

RayenTian commented Jul 24, 2025

Uh oh!

SahilJain314 left a comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

5 participants

Conversation

RayenTian commented Jul 21, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

What does this PR do ?

Issues

Test Result

For remove NRL_SKIP_TIED_WEIGHT_CHECK env variable

Llama-3.2-1B

Qwen2.5-1.5B-Instruct

google/gemma-2-2b-it

For remove model.config.tie_word_embeddings in nemo_rl/models/dtensor/parallelize.py

Qwen/Qwen2.5-1.5B-Instruct

Qwen/Qwen2.5-7B-Instruct

More Custom Test

Additional Notes on Gemma Model Testing

Uh oh!

Uh oh!

Uh oh!

yuki-97 commented Jul 22, 2025

Uh oh!

yuki-97 commented Jul 22, 2025

Uh oh!

SahilJain314 commented Jul 23, 2025

Uh oh!

Uh oh!

RayenTian commented Jul 24, 2025

Uh oh!

SahilJain314 left a comment

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

5 participants

RayenTian commented Jul 21, 2025 •

edited

Loading

For remove `NRL_SKIP_TIED_WEIGHT_CHECK` env variable

For remove `model.config.tie_word_embeddings` in `nemo_rl/models/dtensor/parallelize.py`