PCD Pins the Primary Gradient, Holding 70.9% at 90% Sparsity Where MGDA Falls to 1.2%

Not All Objectives Are Born Equal: Priority-Constrained Descent for Hierarchical Multi-Objective Optimization

Dara Varam, Mohamed I. Alhajri

cs.LG, stat.ML

2026-06-29

PCD pins updates to the primary gradient, distorting it just enough for secondary progress via one τ. ResNet-34/CIFAR-100 holds 70.9% at 90% structured sparsity; MGDA falls to 1.2%.

What problem this solves

Most multi-objective work in deep learning still treats losses as peers. Accuracy is the deliverable. Sparsity, rank, and quantization fidelity are constraints on that deliverable. Weighted sums and gradient-manipulation methods (MGDA, PCGrad, CAGrad, FAMO) stop at Pareto stationarity: the origin sits in the convex hull of the gradients, so no shared descent direction remains.

That stop is the wrong target when the losses are ranked. Gradients can cancel while none of them is individually stationary. The paper calls this a conflict equilibrium. On a 2D toy with a saddle between a local min and a joint min, every symmetric baseline stalls with L2 still nonzero. PCD is the only method in that figure that crosses the barrier and reaches ∇L1 = ∇L2 = 0.

Method

PCD treats the primary gradient as an anchor and the secondaries as half-space constraints. After EMA-normalizing each gradient (one Adam-style second-moment scalar per objective), it finds the Euclidean-closest direction to the primary that still gives every secondary at least a τ-fraction of its own maximum first-order progress, with τ in [0, 1].

The primary coefficient is pinned at one. Secondary coefficients are nonnegative multipliers that vanish when the constraint is already met. If the raw primary direction already clears every margin, the update is ordinary gradient descent on L1.

For two objectives the projection is closed form: add just enough of the secondary gradient to hit the margin with equality. At K=2 on raw gradients this is Dynamic Barrier Gradient Descent (Gong et al., 2021). PCD adds per-objective scale normalization, so one dimensionless τ is comparable across heterogeneous regularizers. Rescaling a secondary by c from 10⁻⁴ to 10⁶ at fixed τ=0.3 moves the operating point by only 0.0075 in original L2 units. The same nominal weight on a weighted sum slides across the whole front.

The guarantees attach to the normalized direction, not to the deployed algorithm. The practical step rescales that direction to the raw primary-gradient norm before handing it to Adam. There is no convergence theorem. When the primary gradient is zero the rescale zeros the update, so a primary-stationary point that is not CMS can be a deployed fixed point. Across 330 tuned-baseline runs of 300 epochs, the smallest primary gradient seen was 4.7×10⁻⁴.

Results

Structured pruning uses cross-entropy as L1 and channel-wise Group Lasso as L2. After training, groups below a threshold are removed with no fine-tuning. Four architectures on CIFAR-10 and CIFAR-100, Adam, 300 epochs, five seeds.

ResNet-34 / CIFAR-100, dense accuracy 73.3%:

MethodAcc. @ 80%Acc. @ 90%
PCD73.0%70.9%
AuxiNash44.4%44.4%
Weighted sum27.3%27.3%
PCGrad / CAGrad7.4% / 7.3%7.4% / 7.3%
MGDA1.2%1.2%

Symmetric methods drive native group sparsity to 99.3% to 100% during training; pruning only reveals a network that is already broken. On DenseNet-121 / CIFAR-10, τ=0.01 keeps 94.0% versus 94.5% dense, with 5.3× fewer FLOPs, 1.8× lower latency, and 9.4× smaller files. At τ=0.05, accuracy is 89.2% with 45.4× fewer FLOPs and 83× smaller files.

Against tuned pruning-specific baselines (log-grid Group Lasso, proximal Group Lasso, Cooper, both DBGD branches), PCD wins 28 of 47 matched comparisons in the band with at least 85% native group sparsity, loses 1, and ties within seed noise on 18. The one loss is real: tuned scalarization reaches 53.8% at 98.1% sparsity on ResNet-34/CIFAR-100, PCD 51.9% at 97.9%. At the same nominal knob τ=β=0.01, PCD accuracy retention across four configs spans 1.09 points; DBGD spans 13.23.

On the three-objective task (cross-entropy, ℓ1, nuclear norm) at τ=0.02, ResNet-34/CIFAR-100 records 71.0% accuracy, 92.6% unstructured sparsity, effective rank 24.5. AuxiNash, the strongest baseline, sits at 48.7% / 98.8% / 184.8. Dominated hypervolume on ResNet-34/CIFAR-10 is 35.4 for PCD versus about 13.1 for the next method.

In the conflict-equilibrium synthetic, PCD drives the primary gradient to about 10⁻¹⁵ and the primary loss to about 10⁻²⁹. MGDA and CAGrad stay pinned to 1.536√(K−1); at K=5 that is 3.07. The QP was feasible in 1785 of 1785 random secondary configurations.

Why it matters

For a task loss plus a structural regularizer, PCD replaces a brittle weight sweep with one geometrically defined scalar. The extra per-step cost is a K×K QP, small next to K backprops. Closed forms cover the K=2 and K=3 cases people actually run.

This is a change of update geometry. The pruning pipeline itself is unchanged. The protocol excludes fine-tuning, per-layer slimming, and stochastic gates, so it does not compete with Network Slimming or magnitude pruning plus fine-tuning as a product path. All networks are CIFAR-scale. Anyone already training with Group Lasso or a nuclear-norm penalty can swap the direction step without changing the losses.

Limitations

The paper states the limits. There is no convergence theorem for the deployed algorithm. Guarantees are about the constructed direction. Adaptive optimizers need not preserve constraint signs: under Adam the secondary-constraint sign survives on 92.3% to 99.8% of steps; SGD with momentum flips it on 47% to 73% of steps, yet both land on the same accuracy-sparsity frontier. For K>2 the secondary polyhedron can be empty; the implementation then drops the constraints and falls back to the primary gradient. Secondaries are assumed convex. Standard autodiff need not return 0 when 0 lies in the subdifferential, so the CMS identification depends on the oracle. On genuine trade-offs with no joint stationary point, the τ>0 direction has no fixed point and iterates need not converge.

The synthetic escape problem is separable. PCGrad also reaches CMS for K≥3 in that geometry, and still loses on every compression target in the structured-pruning study. The comparison omits magnitude pruning with fine-tuning. At 99% compression on MobileNetV2/CIFAR-100, weighted sum scores 1.9% and PCD 1.2%; both have collapsed. τ is most useful in (0, 0.1] on these configs; past 0.1 on ResNet-34/CIFAR-100 accuracy falls sharply. An interpretable knob still has to be swept.

Terms

Source

Related papers

All paper explainers