• Home
  • ニュース
  • NeurIPS 2026に当研究室の論文5本が採録
  • —

    NeurIPS 2026に当研究室の論文5本が採録

    Paper 1
    ■書誌情報

    Yue Xun, Junyu Liu, Qian Niu, Xinyi Wang, Zheng Yuan, Zirui Li, Zequn Zhang, Bowen Zhao, Shujun Wang, Irene Li, Kan Hatakeyama-Sato, Yusuke Iwasawa, Yutaka Matsuo: JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation, Proceedings of the 40th Annual Conference on Neural Information Processing Systems (NeurIPS 2026), Evaluations & Datasets Track, December 2026
    ■概要
    We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF materials released by the Japanese Ministry of Health, Labour and Welfare, JMed48k contains 48,862 exam questions and 20,142 images from 11 national licensing examinations between 2005 and 2025, with visual content annotated under an 8-type taxonomy. From this corpus, we derive JMed48k-Eval, a recent five-year evaluation subset with 12,484 scored questions, including 9,905 text-only questions and 2,579 questions with images. We evaluate 21 proprietary, open-source, and medical-specific models, reporting text-only and with-image performance separately. Because these subsets contain different questions, we further introduce a paired image-removal audit that evaluates questions with images before and after removing visual content to explore four answer-transition states. The audit shows that proprietary and open-source models gain substantially from images, whereas medical-specific systems show limited observable use of visual evidence, with many correct answers persisting after image removal. Even among proprietary models, the net image-removal effect varies sevenfold across professions, from +5.7 points on Physician questions to +39.8 points on Public Health Nurse questions. We release JMed48k to support reproducible, profession-stratified evaluation of vision-language models in medical licensing settings.

    Paper 2
    ■書誌情報

    Bum Jun Kim, Hyeyun Jeong: Rotation-Invariant Vector Normalization for Molecular Force Learning, Proceedings of the 40th Annual Conference on Neural Information Processing Systems (NeurIPS 2026), December 2026
    ■概要
    Rotating a molecule should leave its energy unchanged while rotating each force vector by the same rotation, as required in SO(3)-equivariant molecular learning. However, conventional normalization layers in neural networks are designed for scalar activations; when applied to vector channels through component-wise centering or scaling, they can break this symmetry. We study normalization modules for equivariant molecular networks under force supervision, where invariant scalar channels and covariant vector channels require different geometric treatment. We formalize an SO(3)-equivariance-preserving recipe that computes vector scale statistics from rotation-invariant squared norms and applies the resulting rescaling identically across Cartesian coordinates, which preserves SO(3)-equivariance without vector shifts. Building on this recipe, we introduce grouped vector RMS (GroupRMS), which shares vector RMS estimates across channel groups to reduce estimator variance and stabilize vector scaling in small-batch training. The proposed variant is lightweight and compatible with standard scalar normalization branches. We instantiate these ideas in practical scalar-vector normalization blocks for PaiNN-style backbones and evaluate them across several molecular benchmarks. On the paired three-trajectory subset, the strongest GroupRMS-family variant, GroupNorm-GroupRMS, improves average force MAE from 0.1190 to 0.0383 and average energy MAE from 0.0675 to 0.0424 relative to BatchNorm. These results support the proposed rotation-compatible vector scaling and grouped scale estimation as effective tools for robust molecular force learning.

    Paper 3
    ■書誌情報

    Bum Jun Kim, Gnankan Landry Regis N’guessan: Per-Loss Adapters for Gradient Conflict in Physics-Informed Neural Networks, Proceedings of the 40th Annual Conference on Neural Information Processing Systems (NeurIPS 2026), December 2026
    ■概要
    Physics-informed neural networks (PINNs) train a single neural approximation by minimizing multiple physics- and data-derived losses, but the gradients of these losses often interfere and can stall optimization. Existing remedies typically treat this pathology either through scalar loss balancing or full-parameter-space gradient surgery, leaving it unclear which intervention is most appropriate. We show that PINN gradient conflict is not a uniform failure mode with one universal remedy. Instead, we identify distinct PINN gradient-conflict regimes, each associated with a different intervention class. Persistent directional conflict may require separate loss-indexed parameter subspaces, magnitude imbalance often favors scalar reweighting, and low or transient conflict may require no extra mitigation. To select between scalar reweighting and a lightweight architectural intervention, we propose a diagnostic-first framework. It profiles a 1000-step unmodified PINN run and, when intervention is warranted, uses one low-rank adapter per loss to create explicit loss-indexed parameter subspaces attached to a shared PINN trunk, providing each loss with a direct gradient pathway. Across more than 60 PDE configurations, including forward, inverse, multi-physics, parameter-varying, and high-dimensional problems up to 50D, persistent directional conflict dominates standard forward K=3 benchmarks and a natural K=4 thermoelastic system, where adapters combined with reweighting yield significant improvements. In contrast, K=3 inverse problems and natural K=5 and K=6 multi-physics systems are largely magnitude-dominated and often favor reweighting alone, while full-parameter-space gradient surgery can fail on heterogeneous parameter spaces. A regime-transition theorem and a blockwise neural tangent kernel analysis explain why gains from these separate adapter parameter subspaces arise selectively in persistent-conflict regimes.

    Paper 4
    ■書誌情報

    Bum Jun Kim: Exactness Matters for Physical Rule Enforcement, Proceedings of the 40th Annual Conference on Neural Information Processing Systems (NeurIPS 2026), December 2026
    ■概要
    Autoregressive scientific forecasters often enforce physical or structural constraints by repairing each predicted state before feeding it back into the model. However, it remains unclear when stronger physical rule enforcement becomes reliable and when it becomes a source of distribution shift. We study this question through operator exactness, meaning whether the repair map is the identity on the target manifold and is aligned with the target geometry. We compare raw forecasting, post hoc repair, and in-loop repair across periodic incompressible Navier–Stokes, non-periodic CFDBench flows, and a hierarchical-forecasting support task. In the exact periodic regime, Fourier projection substantially improves rollout accuracy. On the NS-128 benchmark, a strong Raw-FNO has a final-step rollout MSE at horizon 100 of (9.390±6.290)×10−5, and post hoc and in-loop projection reduce it to (1.130±0.165)×10−6 and (5.370±0.113)×10−7. However, once an exact projection is unavailable and only approximate boundary-preserving cleanup is available, the ordering changes. Across cavity, tube, dam, and cylinder flow, stronger Poisson-based cleanup can reduce divergence while worsening rollout error; target-distortion MSE predicts this harm far better than a linear-system residual. Controlled mismatch, screened cleanup, adaptive gating, and external-backbone checks show that the best approximate-regime operating point can be raw or near-identity. Hierarchical forecasting gives the same broader pattern. Exact forecast reconciliation is a stable baseline, whereas blended top-down repair, a validation-tuned interpolation toward historical-proportion top-down reconciliation, is dataset-dependent. Thus, constraint enforcement should be benchmarked by operator–data alignment before enforcement strength. Use in-loop projection when the operator is exact, and validate approximate cleanup strength using rollout metrics otherwise.

    Paper 5
    ■書誌情報

    Shunsuke Yasuki, Soshun Kihara, Masato Taki: The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation, Proceedings of the 40th Annual Conference on Neural Information Processing Systems (NeurIPS 2026), December 2026
    ■概要
    Convolutional neural networks trained on ImageNet are known to exhibit a strong preference for local high-frequency texture, an inductive bias that translates into fragile robustness against distribution shifts in real-world environments. Event cameras, in contrast, record only changes in scene brightness and are therefore well suited to capturing contour information; however, due to the absence of diagnostic benchmarks in the event domain, the inductive bias that event-camera data instills in vision models has remained underexplored. In this work, we use knowledge distillation from the event domain to the RGB domain so as to exploit the rich evaluation toolkit available in the RGB domain and systematically dissect this inductive bias. Our experiments show that distillation from the event domain induces, in the RGB domain, color invariance, shape bias, and robustness to high-frequency noise. We identify the underlying mechanism as the model suppressing its dependence on high-frequency texture while acquiring a stronger dependence on edge-based object shape. This hypothesis is supported by changes in how color and spatial information are processed at the early layers, together with a spectral trade-off in which robustness to the absence of high-frequency components coexists with vulnerability to contamination of the relied-upon frequency bands and to disruption of geometric structure. We further show that this inductive bias differs from existing robustification methods and that it functions as a useful prior for diverse downstream tasks in which shape and contour information contribute alongside other cues.