← Back to homepage

Jiaju Han

Prospective Fall 2027 PhD Applicant

World Models | Embodied Intelligence

Address: Shenzhen, China · Email: hanjiaju05@gmail.com · Website: jiajuhan.com · arXiv: 2607.29445 · 2607.07288 · 2607.06552 · 2607.06485 · 2606.17020 · 2605.07273 · 2605.22273 · 2603.28568 · 2603.27759

Education

China University of Petroleum (Beijing)

B.Eng. in Data Science and Big Data Technology (Second Bachelor's Degree), Sep 2025 - Jun 2027

GPA: 3.77 · Rank: 2/71

China University of Petroleum (Beijing)

B.A. in English, Sep 2021 - Jul 2025

Research Experience

Shenzhen Research Institute of Big Data (SRIBD) and CUHK-Shenzhen

Jun 2026 - Present · Research Assistant · Research Direction: World Models, Embodied Intelligence, and Large Language Models

  • Investigate world-model-based predictive modeling and decision-making for embodied and energy-system scenarios, connecting multimodal representations with planning and control.
  • Explore world models and embodied intelligence for learning environment dynamics, forecasting future states, and supporting planning under uncertainty.
  • Develop an LLM- and GraphRAG-based power-system fault prediction framework that retrieves topology-aware evidence and supports interpretable risk diagnosis.

China University of Petroleum (Beijing)

Sep 2025 - Jun 2026 · Undergraduate Researcher · Research Direction: Multimodal Learning and VLM Robustness

  • Led and co-led multimodal robustness studies across visible, infrared, and remote-sensing VLMs; developed structured attacks and unified evaluation for zero-shot classification, image captioning, and VQA.
  • Built reproducible evaluation workflows spanning CLIP-style encoders and generative VLMs, covering transferability, cross-task semantic drift, retrieval failures, and response analysis.
  • Contributed to 13 conference papers through problem formulation, experimental design, analysis, and writing; five are first-author works, including one accepted at ACM Multimedia 2026.

Publications

When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models

Studies physically plausible wrinkle-induced attention shifts and examines how surface deformations alter visual grounding and transfer across VLM classification, image captioning, and VQA.

ACM Multimedia 2026 (Accepted) · arXiv: 2603.27759

XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs

Proposes an imperceptible X-shaped sparse perturbation that concentrates changes in structured regions and evaluates transferability across multiple VLM architectures and downstream tasks.

Submitted to AAAI 2027 · arXiv: 2603.28568

FusionRS: A Large-Scale RGB-Infrared-Style Remote Sensing Dataset for Cross-Modal Vision-Language Learning

Builds a large-scale RGB-infrared-style remote-sensing dataset and benchmark for cross-modal representation learning, supporting systematic evaluation of CLIP and generative VLM adaptation.

Submitted to the KDD Datasets and Benchmarks Track · arXiv: 2606.17020

MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation

Introduces an infrared remote-sensing vision-language dataset and adapts CLIP and generative VLM backbones with infrared-aware supervision for representation learning and multimodal understanding.

Submitted to ACCV 2026 · arXiv: 2607.06552

InfraQR: Edge-Placed QR-Inspired Structured Patch Attacks on Infrared Vision-Language Models

Introduces an edge-placed QR-inspired structured patch tailored to infrared imagery and evaluates robustness, transferability, and semantic failure modes across infrared VLM tasks.

Submitted to ACCV 2026 · arXiv: 2607.07288

AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models

Models physically motivated thermal-airflow perturbations for infrared remote sensing and evaluates how spatially distributed temperature changes disrupt VLM predictions across tasks.

Submitted to ACCV 2026 · Corresponding author · arXiv: 2607.06485

QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models

Uses QR-structured thermal triggers to induce targeted semantic failures in infrared VLMs and examines targeted effectiveness and transferability across multiple models and tasks.

Submitted to ACCV 2026 · arXiv: 2607.29445

From Clouds to Hallucinations: Atmospheric Retrieval Hijacking in Remote Sensing Vision-Language RAG

Demonstrates how atmospheric retrieval cues can hijack remote-sensing multimodal RAG and evaluates attack effectiveness, cross-dataset transfer, and retrieval-grounding robustness.

Submitted to NeurIPS 2026 · arXiv: 2605.07273

Exposing Vulnerabilities in Visible-Infrared VLMs: A Unified Geometric Adversarial Framework with Cross-Task Transferability

Develops a unified geometric perturbation framework for visible-infrared VLMs and analyzes cross-task transferability across classification, image captioning, and VQA.

Submitted to NeurIPS 2026 · arXiv: 2605.22273

Skills

Languages

  • English: B.A. in English, TEM-4, CET-6
  • Mandarin Chinese: Native

Research

World Models, Embodied Intelligence, Multimodal Learning, Adversarial Attacks, VLM Robustness, RAG, Fine-tuning, Benchmark and Evaluation Design, and LoRA.

Programming

Python, PyTorch, and Hugging Face Transformers.

References

References available upon request.