DVLA-RL++: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning

Wenhao Li 1,2 Xianjing Meng 3 Qiangchang Wang 1 Zhongyi Han 1 Yilong Yin 1 Liqiang Nie 4
1 School of Software, Shandong University 2 Shenzhen Loop Area Institute 3 Shandong University of Finance and Economics 4 Harbin Institute of Technology (Shenzhen)

Journal extension of DVLA-RL (ICLR 2026) IEEE TPAMI ยท under review

TL;DR

Overall framework of DVLA-RL++: semantic banks from labeled supports, intrinsic semantics guiding local and global alignment, competing bank responses selecting support evidence, and paired sampled and reference gate trajectories during policy training.
Richer descriptions do not guarantee better support evidence. DVLA-RL++ decides which support tokens enter a prototype and how strongly semantics influence each visual layer.

Support-evidence contamination

A support image may contain both a discriminative object part and a background pattern that agrees with its class description. Treating both as positive evidence reinforces contextual shortcuts instead of resolving visual ambiguity.

Complementary semantic purification

Intrinsic and nuisance descriptions are generated from labeled supports only and compared in a shared embedding space. An ambiguity-dependent rejection margin and a sparse allocation rule select supported visual tokens, and the rejected mass is assigned to an intrinsic semantic anchor.

Counterfactual RL gating

A stochastic fusion policy is trained against an independently executed reference trajectory on the same episode, with a reward that combines classification utility and nuisance exposure. At inference only the deterministic mean policy runs.

  1. Which support evidence should define a class prototype when both object parts and incidental context agree with the class description?

  2. How much semantic information should enter each visual layer, and how can that decision be learned from a paired, low-variance training signal?

  3. Does deciding both improve recognition and robustness? DVLA-RL++ is the most accurate method in all 18 standard, fine-grained and cross-domain settings, gains 1.4 points on average over DVLA-RL, and shrinks the clean-to-shift accuracy drop from 5.9 to 2.2 points on miniImageNet.

Why does a richer description not guarantee better evidence?

Few-shot learning builds class representations from a handful of labeled supports and compares separate queries with these summaries, so the quality of the evidence entering a prototype directly affects recognition. Our conference framework DVLA-RL combines local attributes and global descriptions with reinforcement-learning gates for hierarchical alignment. Nevertheless, a richer description and a more flexible fusion rule do not guarantee that the visual evidence emphasized by the model belongs to the class rather than to its context.

(a) Intrinsic appearance and incidental context share visual cues, so both agree with support features. (b) Comparing intrinsic and nuisance agreement separates discriminative feathers from contextual reeds, and a rejection margin of 0.15 excludes the ambiguous reed plume despite its positive relative evidence.

A bird among reeds illustrates the difficulty. Striped plumage shares visual patterns with vertical reeds and a reed plume. Both descriptions are factually correct and agree with support features, yet contextual cues become unreliable when the background changes. The reed plume exposes a subtler failure: its intrinsic agreement is high, but its advantage over the nuisance agreement is small, so a positive relative score alone would retain this ambiguous token while the rejection margin excludes it.

Positive image-text agreement does not determine which evidence should define a class. Semantic assistance therefore requires both comparison against a competing explanation and rejection of ambiguous evidence.

DVLA-RL++ keeps dual-level alignment and adds two decisions

An episode contains labeled supports and separate queries from disjoint categories. Class names are available as semantic supervision, whereas query labels are used only for training objectives and evaluation metrics. The method retains the dual-level semantic alignment and scalar attention gate of DVLA-RL, but replaces unconditional support averaging with selective aggregation and learns how much semantic information to inject with a policy whose decisions depend on supports and preceding actions. The additional adaptation is fully specified before a test query is processed and needs no query labels.

A

Semantic banks

Labeled supports generate intrinsic and nuisance phrases at local and global granularity. Generation is cached once per episode and never sees a query.

B

Gated alignment

Intrinsic semantics guide local and global visual alignment through RL-gated attention blocks at two selected layers.

C

Purified prototypes

Competing bank responses select support evidence, which is mixed with an intrinsic anchor and aggregated into class prototypes.

D

Counterfactual policy

Sampled and frozen-reference gates follow separate trajectories, and their paired utilities update the gate policy alone.

RL-gated attention block. Shared intrinsic semantic queries produce image-guided and text-guided contexts, which a support-conditioned scalar gate mixes. The pooled context is broadcast to image tokens, the fused semantic tokens are appended for transformer processing, and only the updated image tokens are retained.

One scalar action per layer

Let \(Z_\ell\) denote image tokens and \(T_\ell\) the projected intrinsic semantic tokens concatenated across all candidate classes. The image-guided and text-guided contexts are mixed by the gate action \(a_\ell\in[0,1]\),

\[ V_\ell=\operatorname{Attn}(T_\ell W_q, Z_\ell W_k^v, Z_\ell W_v^v),\qquad S_\ell=\operatorname{Attn}(T_\ell W_q, T_\ell W_k^t, T_\ell W_v^t),\qquad B_\ell(a_\ell)=a_\ell V_\ell+(1-a_\ell)S_\ell, \]

and the pooled context is broadcast to the image tokens before the next block, \(Z_{\ell+1}=\big[F_{\phi,\ell}([Z_\ell+\mathbf{1}\,\mathrm{mean}(B_\ell);\,B_\ell])\big]_{\mathrm{img}}\). The same semantic tokens subsequently process every query, so no query class is needed to choose a description, and the support-derived schedule keeps query representations compatible with the prototypes. Classification uses cosine logits between a normalized query feature and the class prototypes with temperature \(\tau_c\).

Scoring projections are calibrated on labeled base supports before policy training and then frozen, so the policy cannot lower its nuisance cost by rotating the scoring geometry.

Complementary semantic purification

We distinguish semantic levels from evidence roles. Local attributes and global descriptions specify the granularity of alignment, whereas intrinsic and nuisance banks specify whether a description explains the class or its incidental context. "Striped plumage" and "vertical reeds" describe different evidence within the same image, and a nuisance phrase is neither a negated sentence nor a competing class label. Nuisance phrases are used only for scoring and policy state construction.

Projected support tokens compete against intrinsic and nuisance banks. Signed evidence and an overlap-dependent margin determine sparse weights on the original visual features. Each support representation combines retained visual evidence with an intrinsic anchor carrying the remaining mass, and class-wise averaging followed by normalization yields the prototype.

Compare, reject, allocate, anchor

Bank responses use size-normalized log-sum-exp aggregation so that bank size alone cannot increase evidence, and the signed evidence of a token is the difference between its intrinsic and nuisance responses, \(d_{c,k,i}=u^{+}_{c,k,i}-u^{-}_{c,k,i}\). The rejection margin grows with the ambiguity between the two explanations,

\[ o_c=\tfrac{1}{2}\big(1+\max_{r,t}(e^{+}_{c,r})^{\top}e^{-}_{c,t}\big),\qquad \xi_c=\xi_0+\kappa_o\,o_c, \]

and the allocator projects the margin-corrected evidence onto the capped simplex,

\[ \mathbf{a}_{c,k}=\underset{a\ge 0,\;\mathbf{1}^{\top}a\le 1}{\arg\min}\;\tfrac{1}{2}\Big\|a-\frac{\mathbf{d}_{c,k}-\xi_c\mathbf{1}}{\tau_s}\Big\|_2^2,\qquad \rho_{c,k}=\mathbf{1}^{\top}\mathbf{a}_{c,k}. \]

Its threshold solution implies that retained tokens necessarily clear the rejection margin, and unused mass is neutral rather than redistributed among rejected tokens. With a mixing coefficient \(\lambda\) and the normalized intrinsic anchor \(r_c\), the raw support representation is \(w_{c,k}=\lambda\sum_i a_{c,k,i}z_{c,k,i}+(1-\lambda\rho_{c,k})\,r_c\), so a support with no retained token contributes exactly its anchor. The averaged representation is normalized into the prototype \(p_c\).

0.74 → 0.16
Histogram overlap between intrinsic and nuisance signed margins on miniImageNet (DVLA-RL → DVLA-RL++)
0.75 → 0.12
Histogram overlap on CUB
24%
Token retention at the default setting
94%
Intrinsic precision among retained tokens
0.08
Neutral mass assigned to the intrinsic anchor
Signed-margin distributions of intrinsic and nuisance evidence on miniImageNet and CUB for DVLA-RL (top) and DVLA-RL++ (bottom). OVL denotes histogram overlap, and the dotted line marks zero. Nuisance evidence is pushed below the margin while intrinsic evidence remains above it.

Prototype stability. If nuisance scores have mean at most \(-\Delta\) with conditionally sub-Gaussian error of scale \(\sigma\), the expected retained nuisance mass decays as \(\exp[-(\Delta+\xi_c)^2/2\sigma^2]\), and the normalized prototype error is at most twice the raw bound. No independence among tokens is required.

Counterfactual reinforcement-learning gating

Training alternates representation learning and policy learning. A policy batch is collected with the representation, scorer, text encoder and reference policy frozen. The sampled policy and the frozen DVLA-RL reference follow separate trajectories from the same initial inputs, and their final prototypes and query predictions define a paired reward. A policy update changes only gate parameters, a separate classification step updates the representation under stopped actions, and stale trajectories are discarded whenever the representation changes.

Sampled actions and frozen-reference mean actions generate separate state trajectories under shared inputs. Utilities combine recognition and nuisance costs, with a deviation penalty for the sampled branch. The utility gap and a state baseline define the advantage for policy-only CF-PPO updates, during which the encoders remain fixed.

Paired returns on the same episode

The policy observes pooled support features, the bank summaries, rejection statistics and the layer index, and keeps the scalar action space of the original gate through a mean-bounded Beta distribution, \(a_\ell\sim\mathrm{Beta}\big(\kappa_g p_\theta(\mathcal{H}_\ell),\,\kappa_g[1-p_\theta(\mathcal{H}_\ell)]\big)\). The reference executes its own mean actions \(a^0_\ell\) and never receives states produced by sampled actions. Nuisance exposure is measured on post-action support tokens, \(C(\mathbf{a})=\frac{1}{LNK}\sum_{\ell,c,k}\frac{1}{M_\ell}\sum_i[\xi_c-d^{\mathrm{post}}_{\ell,c,k,i}(\mathbf{a})]_+\), and the paired reward and advantage are

\[ R(\mathbf{a})=-\mathcal{L}_{\mathrm{cls}}(\mathbf{a})-\gamma\,C(\mathbf{a})-\frac{\eta}{L}\sum_\ell (a_\ell-a^0_\ell)^2,\qquad R_0=-\mathcal{L}_{\mathrm{cls}}(\mathbf{a}^0)-\gamma\,C(\mathbf{a}^0),\qquad A_\ell=\operatorname{sg}\big[R(\mathbf{a})-R_0-b_\omega(\mathcal{H}_\ell)\big]. \]

With action-independent reference randomness, differentiable positive policy densities and fixed environment parameters, the score-function estimator \(\sum_\ell A_\ell\nabla_\theta\log\pi_\theta(a_\ell\mid\mathcal{H}_\ell)\) is unbiased for the gradient of the expected reward (Proposition 1). Repeated minibatch updates use the clipped surrogate \(\mathcal{L}_{\mathrm{CF\text{-}PPO}}=-\mathbb{E}\sum_\ell\min(r_\ell A_\ell,\,\mathrm{clip}(r_\ell,1-\epsilon_p,1+\epsilon_p)A_\ell)\), and the deviation of its gradient from the importance-weighted gradient is bounded in terms of the active clipped samples.

0.60 vs 1.25
Gradient-covariance trace of paired returns with a state baseline versus raw returns, relative to the state-baseline estimator (40% reduction)
24% → 9%
Frequency of sampled trajectories with higher nuisance cost than the reference
22% → 15%
Active PPO clipping after policy updates
1 forward pass
Per query at inference: banks, prototypes and the mean gate schedule are cached per episode, with no reference rollout or value evaluation
Policy diagnostics on miniImageNet. (a) Gradient-covariance trace relative to the state-baseline estimator, where Paired denotes paired returns with a state baseline. (b) Nuisance-cost increase and active clipping frequencies.

Detaching a tensor is an implementation operation, whereas conditional action independence is a statistical condition. The separate reference rollout provides the latter, which stopping gradients through an action-dependent reference would not.

State-of-the-art accuracy on standard, fine-grained and cross-domain benchmarks

The visual backbone is Visformer-Tiny with 224×224 inputs, the frozen text encoder is the CLIP ViT-B/16 text encoder, and Qwen2.5-VL-32B generates attribute-level and description-level text from labeled supports. For each setting, 2,000 5-way tasks are sampled from novel classes with 15 queries per class, and we report mean accuracy with its 95% confidence interval. Competitor results are quoted from their original reports.

MethodVenueminiImageNettieredImageNetCIFAR-FS
1-shot5-shot1-shot5-shot1-shot5-shot
ProtoNetNeurIPS 201762.39 ±0.2180.53 ±0.1468.23 ±0.2384.03 ±0.1672.20 ±0.7083.50 ±0.50
AM3NeurIPS 201965.30 ±0.4978.10 ±0.3669.08 ±0.4782.58 ±0.31--
DeepEMDCVPR 202065.91 ±0.8282.41 ±0.5671.16 ±0.8786.03 ±0.58--
SUNECCV 202267.80 ±0.4583.25 ±0.3072.99 ±0.5086.74 ±0.33--
FewTURENeurIPS 202268.02 ±0.8884.51 ±0.5372.96 ±0.9286.43 ±0.6772.80 ±0.8886.14 ±0.64
SVAECVPR 202274.84 ±0.2383.28 ±0.4076.98 ±0.6585.77 ±0.50--
Meta-AdaMNeurIPS 202359.89 ±0.4977.92 ±0.4365.31 ±0.4885.24 ±0.35--
ProtoDiffNeurIPS 202366.63 ±0.2183.48 ±0.1572.95 ±0.2485.15 ±0.18--
CPEAICCV 202371.97 ±0.6587.06 ±0.3876.93 ±0.7090.12 ±0.4577.82 ±0.6688.98 ±0.45
SPCVPR 202372.31 ±0.4083.42 ±0.3078.03 ±0.4688.55 ±0.3282.18 ±0.4088.24 ±0.32
ALFATPAMI 202466.61 ±0.2881.43 ±0.2570.29 ±0.4086.17 ±0.3576.32 ±0.4386.73 ±0.31
LastShotTPAMI 202467.35 ±0.2082.58 ±0.1472.43 ±0.2385.82 ±0.1676.76 ±0.2187.49 ±0.12
MetaKernelTPAMI 202466.9 ±0.381.3 ±0.2--76.1 ±0.486.9 ±0.5
SIFTIJCV 202477.31 ±0.6786.95 ±0.5377.86 ±0.7789.89 ±0.52--
SemFewCVPR 202478.94 ±0.6686.49 ±0.5082.37 ±0.7789.89 ±0.5284.34 ±0.6789.11 ±0.54
ECERAAAI 202581.14 ±0.15-81.81 ±0.51-86.01 ±0.35-
CPLTPAMI 202572.82 ±0.6187.93 ±0.3778.05 ±0.7090.89 ±0.4478.82 ±0.6589.98 ±0.44
FRN+IAMTPAMI 202566.96 ±0.1983.19 ±0.1371.85 ±0.2286.55 ±0.15--
DVLA-RLICLR 202681.69 ±0.3688.25 ±0.2883.02 ±0.4391.71 ±0.2987.18 ±0.4090.59 ±0.31
DVLA-RL++Ours83.41 ±0.3489.63 ±0.2684.57 ±0.4192.89 ±0.2788.70 ±0.3891.75 ±0.29

5-way accuracy (%) with 95% confidence intervals on miniImageNet, tieredImageNet and CIFAR-FS.

DVLA-RL++ achieves the best accuracy in every setting. It exceeds the strongest compared method by 1.54 percentage points on average over the twelve in-domain settings and improves on DVLA-RL by 1.03 points on average across the six cross-domain settings, so purified support evidence transfers across domains with different visual statistics.

Each module improves accuracy, and both together improve robustness

Ablations use the 5-way 1-shot setting with matched semantic caches and training budgets, five training seeds and 2,000 evaluation episodes per seed. Each variant reports clean accuracy and the accuracy on matched queries after a context intervention that holds the foreground object and its label fixed while replacing the background or an incidental co-occurring object.

Component controls: clean versus context-shifted queries
CleanShift

CSP denotes complementary semantic purification and CFG counterfactual gating. Purification alone retains the conference gate, whereas gating alone retains uniform visual support prototypes. Hover a bar for its value.

VariantminiImageNetCUB
CleanShiftCleanShift
DVLA-RL81.775.891.985.0
Uniform allocation82.277.692.887.2
Positive semantics only82.779.093.589.1
No neutral option83.079.793.890.0
Static gate82.879.893.690.0
Fixed margin83.180.194.090.6
No intrinsic fallback82.979.693.789.8
State baseline only83.080.494.090.9
No nuisance cost83.180.194.090.5
DVLA-RL++83.481.294.392.0

Internal purification and policy controls on 5-way 1-shot accuracy (%). Uniform allocation replaces the selection weights by 1/M, positive-only semantics removes the nuisance bank, the no-neutral control forces the allocation mass to one, and the static control learns a support-independent scalar gate.

Every control reduces shifted accuracy more than clean accuracy, which confirms that signed evidence, neutral allocation, ambiguity-aware rejection and the counterfactual training signal each contribute to robustness against contextual change.

Parameter sensitivity to (a) the allocation temperature τs, (b) the visual-anchor coefficient λ, and (c) the nuisance-cost weight γ. For τs ∈ {0.02, 0.05, 0.1, 0.2, 0.5, 1.0} the miniImageNet accuracies are 79.8, 81.6, 83.1, 83.4, 83.0 and 82.2%. Accuracy stays within 0.3 points of the default for λ in [0.5, 0.9] and within 0.2 points for γ in [0.1, 0.75].
t-SNE projections of query features on miniImageNet, CIFAR-FS, Stanford Dogs and Places. Baseline features overlap because of contextual variation, whereas DVLA-RL++ features form compact class neighborhoods.
Attention of the baseline, purification alone, gating alone and DVLA-RL++ on nine images covering co-occurrence, occlusion, clutter and camouflage. Purification controls which support evidence enters the prototype, while gating controls how semantics influence successive visual layers.

Citation

@article{li2026dvlarlpp,
  title   = {DVLA-RL++: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning},
  author  = {Li, Wenhao and Meng, Xianjing and Wang, Qiangchang and Han, Zhongyi and Yin, Yilong and Nie, Liqiang},
  journal = {arXiv preprint arXiv:2610.12095},
  year    = {2026}
}
@inproceedings{li2026dvlarl,
  title     = {DVLA-RL: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning},
  author    = {Li, Wenhao and Meng, Xianjing and Wang, Qiangchang and Han, Zhongyi and Wu, Zhibin and Yin, Yilong},
  booktitle = {International Conference on Learning Representations (ICLR)},
  year      = {2026}
}