Skip to content

[tunix] Add weighted prepared diffusion OPD objective#1747

Open
ethannnnnn wants to merge 5 commits into
google:mainfrom
ethannnnnn:block-diffusion-tunix-pr5-opd
Open

[tunix] Add weighted prepared diffusion OPD objective#1747
ethannnnnn wants to merge 5 commits into
google:mainfrom
ethannnnnn:block-diffusion-tunix-pr5-opd

Conversation

@ethannnnnn

Copy link
Copy Markdown

Motivation

Prepared diffusion rollouts need a reusable objective that combines teacher distributions, optional hard targets, explicit fractional token weights, and correct trainer/checkpoint behavior.

Scope

  • Add forward KL(teacher || student) with temperature-squared scaling.
  • Add optional target-aligned hard cross entropy.
  • Support fractional weights, inactive-token sanitization, and teacher stop-gradient.
  • Wire the objective through PeftTrainer and preserve custom semantic metadata on periodic and final checkpoints.

Design

configure_prepared_diffusion_opd consumes an externally prepared fresh batch. Student logits come from the canonical target-aligned scorer; teacher logits are treated as constants. KL and hard CE share the explicit weights and denominator-aware LossOutput reduction.

Trainer metadata hooks default to an empty mapping, while model-aware integrations can record the semantic identity required to reject incompatible resume attempts.

Compatibility

Existing trainers and distillation objectives remain unchanged unless this prepared objective is selected. The metadata hook is empty by default for existing users.

Extensibility

Model integrations own rollout and alignment, so the same Tunix objective can support additional diffusion architectures. A future top-k teacher representation can be introduced behind the prepared-batch boundary.

Tests

  • 11 focused OPD tests.
  • Cumulative suite: 34 diffusion contract/SFT/prepared-batch/OPD tests plus 8 weighted-trainer regressions.
  • Checkpoint tests cover custom metadata on periodic and forced final saves in classic PeftTrainer and PeftTrainerV2.
  • Formatter, import-order, type, lint, compilation, and diff checks passed for the scoped stack.

Known limitations

Tunix neither performs nor verifies on-policy generation. Dense teacher logits are required, and external artifact references must be pinned by the caller if reproducible resume is required.

Stack

Depends on the preceding upstream PR: #1746

Tunix block-diffusion design document

Define a target-aligned, batch-major diffusion batch contract and typed adapter/scorer protocols without depending on MaxText or a specific training algorithm.

Validate shapes and dtypes at construction and scoring boundaries, while preserving JAX pytree, JIT, and sharding compatibility.

Tests: 10 diffusion contract tests; pyink/isort; pylint; pyrefly; py_compile.
Accumulate LossOutput gradients as unreduced sums and normalize once by the
total denominator across microbatches. Preserve denominator-one behavior for
scalar losses and return zero gradients when every weight is zero.

Select auxiliary-metric reducers by value type in training and evaluation:
globally combine weighted metrics while averaging ordinary scalar metrics.
Reject per-key type changes across microbatches and preserve consistent
epsilon and minimum-denominator bounds during global reduction.

Preserve the dtype selected by each Optax optimizer-state initializer across
conditional update and skip branches. This keeps explicit bf16 moments in
bf16, retains explicit fp32 moments, and prevents Flax NNX branch-type
mismatches without special-casing a particular accumulation count.

Tests cover weighted and fractional denominators, zero-weight batches, mixed
weighted/plain train and eval metrics, reducer invariants, and a real
PeftTrainer + nnx.jit matrix over direct/injected AdamW and gradient
accumulation counts 1 and 2. The complete PeftTrainer suite passes 56 tests;
the cumulative focused validation passes 119 tests with six optional engine
tests deselected. Ruff and git diff checks pass.
Provide a typed PeftTrainer adapter for canonical diffusion batches and target-aligned score functions. Compute weighted float32 cross entropy without autoregressive shifting, sanitize inactive targets, and preserve zero-weight numerical safety.

Tests: 17 diffusion contract and SFT tests; 6 focused weighted-gradient tests; pyink, isort, pyrefly, pylint, py_compile, and diff checks.
Define a framework-neutral external-teacher batch contract for freshly prepared student rollouts. Validate the canonical student batch and target-aligned teacher logits without owning model rollout, corruption, or checkpoint behavior.

Tests: 3 focused batch-contract tests; included in the 34-test diffusion contract/SFT/OPD suite.
Add forward teacher-to-student KL with temperature scaling, optional target-aligned hard CE, fractional token weights, teacher stop-gradient, inactive-token sanitization, and PeftTrainer wiring for externally prepared fresh rollouts.

Tests: 11 focused OPD tests; 34 combined diffusion contract/SFT/OPD tests and 8 weighted trainer regressions passed.
@google-cla

google-cla Bot commented Jul 23, 2026

Copy link
Copy Markdown

Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA).

View this failed invocation of the CLA check for more information.

For the most up to date status, view the checks section at the bottom of the pull request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants