LLM Post-Training for Time-Series Anomaly Detection
Controllable trajectory synthesis, Agentic RL, multi-turn tool use, and VLM active perception for intent alignment and fine-grained localization.
View Work →I am a Ph.D. student in the Department of Computer Science and Technology at Tsinghua University, advised by Prof. Dan Pei in NetMan Lab. My research focuses on LLM post-training, Agentic RL, multimodal LLM alignment and reasoning, and their applications to AIOps.
My recent work focuses on LLM post-training for time-series anomaly detection, with an emphasis on user-intent alignment, memory retrieval, multi-turn tool use, and VLM-based active perception. Along the way, I identified off-policy effects caused by asynchronous policy updates and trajectory generation, which led me to study off-policy optimization for LLM post-training and long-horizon Agentic RL, particularly the impact of policy lag and historical trajectory drift.
I received my B.E. degree in Computer Science and Technology from Tsinghua University in 2023. I have held algorithm research internships at ByteDance, Huawei, ZTE, eBay, and RealAI, producing two first-author ICML 2026 papers and two additional first-author submissions.
Current work spans LLM post-training, Agentic RL, multimodal reasoning, and their applications to AIOps.
Controllable trajectory synthesis, Agentic RL, multi-turn tool use, and VLM active perception for intent alignment and fine-grained localization.
View Work →Off-policy learning in LLM RL and long-horizon agents under policy lag, stale trajectory reuse, and accumulated decision drift.
Current researchNative-modality alignment, structured reasoning, and efficient serving for time-series data in LLMs.
View Work →High-quality evaluation data, comprehensive benchmarks, and automatic scoring methods for domain-specific LLMs.
View Work →Recent research updates and publications.
For the full list, see my Google Scholar profile.
Research FocusLLM Post-Training for Time-Series Anomaly Detection
Builds a multi-step agentic decision process across intent parsing, autonomous memory retrieval, and tool execution, then injects intent-following and decision capabilities through Agentic RL; deployed in ByteDance's development environment.
Research FocusLLM Post-Training for Time-Series Anomaly Detection
Builds a multi-turn inspection agent on Qwen3.5-9B that screens candidate regions globally, gathers high-resolution local evidence in parallel, refines boundaries across turns, and stops adaptively; trained through controllable data synthesis and two-stage SFT + GRPO.

Research FocusTime-Series Foundation Models
The first numerical-sequence forecasting model robust to anomalies in recent observations; it analyzes shortcut reliance on recent data, uses random-sampling-based global-context modeling, and combines learnable period extraction with two-stage training.

Research FocusTime-Series Foundation Models
A training-free inference-time enhancement framework that combines downsampling and spline interpolation to expose longer contexts, improving forecasting accuracy by 19% while reducing peak memory by 6.4x and inference time by 16.9x.

Research FocusLLM Evaluation & AIOps
Presents MicroScope, a practical root-cause localization algorithm for complex internal systems that models causal relationships among metrics in latent space and identifies root causes through intervention.

Research FocusMultimodal LLM Alignment & Reasoning
One of the first LLM-based foundation models for fault-diagnosis reasoning, organizing multimodal semantic features, anomaly fluctuations, and propagation relationships into structured reasoning chains and improving generalization with GRPO + RLVR.

Research FocusMultimodal LLM Alignment & Reasoning
The first multimodal LLM that incorporates time-series data as a native new modality; enabled multimodal alignment through SFT, integrated the modality into vLLM for efficient inference, and reached 470+ GitHub stars, 120+ citations, and 20K+ model downloads as of Aug. 2026.

Research FocusLLM Evaluation & Benchmarking
A comprehensive AIOps benchmark with 9,000+ QA pairs across 8 subdomains and multiple evaluation paradigms, plus FAE-Score, a QA scoring metric aligned with expert evaluation.

Research FocusLLM Evaluation & Benchmarking
An automatic evaluation-dataset generation framework based on hierarchical operations documents, combining seed-data synthesis, hierarchical document retrieval, and document-consistency quality verification; deployed in Huawei's practical development environment.

Research FocusLLM Evaluation & Benchmarking
Automates technical-support QA evaluation with key-term matching, step-order verification, and completeness checks, outperforming the previous state of the art by 7.6% AUC.

Research FocusRobust AI Systems
Develops adversarial textured 3D meshes that attack open-source face-recognition models across architectures and bypass real-world defense systems, including phone unlock and smart-door-lock systems.

Research FocusRobust AI Systems
Optimizes face patches in the 3D StyleGAN latent space to achieve robust physical adversarial attacks against face-recognition systems.
Industry research experience across LLM post-training, multimodal reasoning, AIOps, numerical forecasting, and fault diagnosis.
Research Intern. Built multimodal LLM post-training and evaluation pipelines for time-series reasoning, covering controllable data synthesis, multi-turn tool-use trajectories, SFT, Agentic RL, trajectory evaluation, RLVR, and efficient serving. Extended LLaMA-Factory, VeRL, TRL GRPOTrainer, and vLLM for native time-series inputs; work spans ChatTS, FoundRoot, IntentAD, and G2AD.
Research Intern. Built an end-to-end AIOps workflow for domain corpus and QA synthesis, LLM post-training, and automatic evaluation, improving domain QA accuracy by more than 80% over the chat base model. Co-authored OpsEval (FSE 2025) and TechSupportEval (IJCNN 2025); the public leaderboard has exceeded 2 million visits.
Research Intern. Studied robustness and inference efficiency for time-series foundation models, addressing recent-observation shortcuts and long-context limitations. Produced two first-author papers, TameR and SPRINT (ICML 2026).
Research Intern. Studied multisource spatio-temporal foundation models for optical transport networks and automatic evaluation-data generation from hierarchical documents; contributed to Eagle (FSE 2026, CCF A) and supported engineering deployment.
Research Intern. Developed root-cause localization for complex internal systems by modeling causal relationships among metrics in latent space and identifying root causes through intervention. The algorithm was engineered for practical use and produced MicroScope (KDD 2024, CCF A).
Research Intern. Studied physical-world 3D adversarial attacks for face-recognition model security, including adversarial textured meshes and 3D latent-space patch optimization. Co-authored AT3D (CVPR 2023, CCF A, Highlight) and Face3DAdv (IJCV 2024, CCF A).
Selected awards and recognitions.