Task arithmetic provides an efficient, training-free way to edit pre-trained models, yet lacks a fundamental theoretical explanation for its success. The existing concept of "weight disentanglement" describes the ideal outcome of non-interfering task composition but does not reveal its underlying cause. Crucially, what intrinsic properties of the pre-trained model ($\theta_0$) or the task vectors ($\tau_t$) enable this disentanglement remains underexplored. In this paper, we introduce Task-Feature Specialization (TFS), a model's ability to allocate distinct internal features to different tasks, as the fundamental principle. We first prove that TFS is a sufficient condition for weight disentanglement. More importantly, we find that TFS also gives rise to an observable geometric consequence: weight vector orthogonality. This positions TFS as the common cause for both the desired functional outcome (disentanglement) and a measurable geometric property (orthogonality). This relationship provides the key insight for our method: since the abstract TFS property is intractable to enforce directly, we can instead promote weight disentanglement by shaping its concrete geometric consequence, orthogonality. Therefore, we propose OrthoReg, a simple and effective regularization method that actively enforces an internal orthogonal structure on weight updates ($\Delta W$) that constitute $\tau_t$ during fine-tuning. And we theoretically prove that OrthoReg promotes disentanglement. Extensive experiments demonstrate that OrthoReg consistently and significantly enhances the performance of various task arithmetic methods.
✨ Key Contributions
📐 Theory: We identify TFS as a sufficient condition for weight disentanglement, and WVO as its geometric consequence, providing the first principled explanation for task arithmetic.
🔧 Method (OrthoReg): A simple regularization term added to the fine-tuning loss that enforces column-wise orthogonality on ΔW, for which we prove theoretical efficacy.
🔗 Connection to TTA: We show that OrthoReg and Tangent Task Arithmetic (TTA) share the same underlying mechanism (i.e. inter-task vector orthogonality), but OrthoReg achieves this more efficiently.
📊 Experiments: Consistent and significant improvements over Non-linear FT, TTA, ATT-FT, LoRA-ATT across ViT-B-32, ViT-B-16, and ViT-L-14.
The OrthoReg Loss on LoRA-ATT
The OrthoReg loss is applied to the equivalent dense weight update implied by each LoRA module:
To evaluate the OrthoReg variant, replace --finetuning-mode loraatt with --finetuning-mode loraatt_ortho and add --ortho-lambda 10.0 (or 1.0 for ViT-L-14).
Argument reference
Argument
Value for these checkpoints
--seed
1993
--lr
1e-3
--lora-rank
8
--lora-alpha
8.0
--ortho-lambda
0 for loraatt; 10.0 for B-32/B-16, 1.0 for L-14 with loraatt_ortho
For dataset preparation, follow the instructions in the TTA repository.
📝 Citation
If you find this work useful, please cite:
bibtex
1@inproceedings{liu2026orthoreg,
2 title = {Understanding and Enforcing Weight Disentanglement in Task Arithmetic},
3 author = {Liu, Shangge and Yin, Yuehan and Wang, Lei and Fan, Qi and
4 Shi, Yinghuan and Li, Wenbin and Gao, Yang and Tao, Dacheng},
5 booktitle = {CVPR},
6 year = {2026}
7}