Ph.D. Student • Tsinghua University

Longlong Xu

I am a Ph.D. student in the Department of Computer Science and Technology at Tsinghua University, advised by Prof. Dan Pei in NetMan Lab. My research focuses on LLM post-training, Agentic RL, multimodal LLM alignment and reasoning, and their applications to AIOps.

My recent work focuses on LLM post-training for time-series anomaly detection, with an emphasis on user-intent alignment, memory retrieval, multi-turn tool use, and VLM-based active perception. Along the way, I identified off-policy effects caused by asynchronous policy updates and trajectory generation, which led me to study off-policy optimization for LLM post-training and long-horizon Agentic RL, particularly the impact of policy lag and historical trajectory drift.

I received my B.E. degree in Computer Science and Technology from Tsinghua University in 2023. I have held algorithm research internships at ByteDance, Huawei, ZTE, eBay, and RealAI, producing two first-author ICML 2026 papers and two additional first-author submissions.

2 first-author ICML papers LLM post-training & Agentic RL 120+ ChatTS citations
Longlong Xu profile photo.

Research Focus

Current work spans LLM post-training, Agentic RL, multimodal reasoning, and their applications to AIOps.

LLM Post-Training for Time-Series Anomaly Detection

Controllable trajectory synthesis, Agentic RL, multi-turn tool use, and VLM active perception for intent alignment and fine-grained localization.

View Work →

Off-Policy Optimization for LLM & Agentic RL

Off-policy learning in LLM RL and long-horizon agents under policy lag, stale trajectory reuse, and accumulated decision drift.

Current research

Multimodal LLM Alignment & Reasoning

Native-modality alignment, structured reasoning, and efficient serving for time-series data in LLMs.

View Work →

LLM Evaluation & AIOps

High-quality evaluation data, comprehensive benchmarks, and automatic scoring methods for domain-specific LLMs.

View Work →

Latest News

Recent research updates and publications.

May 2026 TameR and SPRINT were accepted to ICML 2026.
Apr. 2026 Eagle, a benchmark question generation framework based on operations documents, was accepted to FSE 2026.
Jan. 2026 FoundRoot was accepted to ICSE 2026.
2025 ChatTS was accepted to VLDB 2025, and OpsEval was accepted to FSE 2025.

Selected Work & Publications

For the full list, see my Google Scholar profile.

View all publications →
Intent-Aligned
Agent
SubmittedFirst author

Research FocusLLM Post-Training for Time-Series Anomaly Detection

IntentAD: User-Intent-Aligned KPI Anomaly Detection Agent

Longlong Xu, et al.

Builds a multi-step agentic decision process across intent parsing, autonomous memory retrieval, and tool execution, then injects intent-following and decision capabilities through Agentic RL; deployed in ByteDance's development environment.

Active
Perception
SubmittedFirst author

Research FocusLLM Post-Training for Time-Series Anomaly Detection

G2AD: VLM-Based Active Perception for Time-Series Anomaly Detection

Longlong Xu, et al.

Builds a multi-turn inspection agent on Qwen3.5-9B that screens candidate regions globally, gathers high-resolution local evidence in parallel, refines boundaries across turns, and stops adaptively; trained through controllable data synthesis and two-stage SFT + GRPO.

TameR paper figure.
ICML 2026 · CCF AFirst author

Research FocusTime-Series Foundation Models

Taming the Recent-Data Bias: Towards Robust Time Series Forecasting with Global Context

Longlong Xu, Zeyan Li, Xiao He, Zhaoyang Yu, Changhua Pei, Zhe Xie, Zijun Dou, Tieying Zhang, Dan Pei

The first numerical-sequence forecasting model robust to anomalies in recent observations; it analyzes shortcut reliance on recent data, uses random-sampling-based global-context modeling, and combines learnable period extraction with two-stage training.

SPRINT paper figure.
ICML 2026 · CCF AFirst author

Research FocusTime-Series Foundation Models

See More, Forecast Better and Faster: Enhancing Time Series Foundation Models via Inference-Time Plug-and-Play Downsampling

Longlong Xu, Zeyan Li, Xiao He, Zhaoyang Yu, Dazhong Wen, Mingze Sun, Changhua Pei, Dan Pei

A training-free inference-time enhancement framework that combines downsampling and spline interpolation to expose longer contexts, improving forecasting accuracy by 19% while reducing peak memory by 6.4x and inference time by 16.9x.

MicroScope root-cause localization framework figure.
KDD 2024 · CCF ACo-author · 3rd student author

Research FocusLLM Evaluation & AIOps

Microservice Root Cause Analysis With Limited Observability Through Intervention Recognition in the Latent Space

Zhe Xie, Shenglin Zhang, Yitong Geng, Yao Zhang, Minghua Ma, Xiaohui Nie, Zhenhe Yao, Longlong Xu, Yongqian Sun, Wentao Li, Dan Pei

Presents MicroScope, a practical root-cause localization algorithm for complex internal systems that models causal relationships among metrics in latent space and identifies root causes through intervention.

FoundRoot structured deep thinking figure.
ICSE 2026 · CCF ACo-author · 2nd student author

Research FocusMultimodal LLM Alignment & Reasoning

FoundRoot: Towards Foundation Model for Root Cause Analysis via Structured Deep Thinking

Zhe Xie, Zeyan Li, Xiao He, Shenglin Zhang, Longlong Xu, Yuzhuo Yang, Tieying Zhang, Jianjun Chen, Rui Shi, Dan Pei

One of the first LLM-based foundation models for fault-diagnosis reasoning, organizing multimodal semantic features, anomaly fluctuations, and propagation relationships into structured reasoning chains and improving generalization with GRPO + RLVR.

ChatTS time-series multimodal LLM figure.
VLDB 2025 · CCF A · 120+ citations4th author · 2nd student author

Research FocusMultimodal LLM Alignment & Reasoning

ChatTS: Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and Reasoning

Zhe Xie, Zeyan Li, Xiao He, Longlong Xu, Xidao Wen, Tieying Zhang, Jianjun Chen, Rui Shi, Dan Pei

The first multimodal LLM that incorporates time-series data as a native new modality; enabled multimodal alignment through SFT, integrated the modality into vLLM for efficient inference, and reached 470+ GitHub stars, 120+ citations, and 20K+ model downloads as of Aug. 2026.

OpsEval framework figure.
FSE 2025 · CCF A3rd author · 2nd student author

Research FocusLLM Evaluation & Benchmarking

OpsEval: A Comprehensive Benchmark Suite for Evaluating Large Language Models' Capability in IT Operations Domain

Yuhe Liu, Changhua Pei, Longlong Xu, Bohan Chen, Mingze Sun, Zhirui Zhang, Yongqian Sun, Shenglin Zhang, Kun Wang, Haiming Zhang, Jianhui Li, Gaogang Xie, Xidao Wen, Xiaohui Nie, Minghua Ma, Dan Pei

A comprehensive AIOps benchmark with 9,000+ QA pairs across 8 subdomains and multiple evaluation paradigms, plus FAE-Score, a QA scoring metric aligned with expert evaluation.

Eagle benchmark generation framework figure.
FSE 2026 · CCF A4th author · 3rd student author

Research FocusLLM Evaluation & Benchmarking

Eagle: Leveraging Operations Documents for Comprehensive Benchmark Question Generation

Yuhe Liu, Changhua Pei, Hang Wang, Longlong Xu, Xiaogang Dong, Zhen Feng, Li Zheng, Kehang Ji, Dan Pei

An automatic evaluation-dataset generation framework based on hierarchical operations documents, combining seed-data synthesis, hierarchical document retrieval, and document-consistency quality verification; deployed in Huawei's practical development environment.

TechSupportEval comparison figure.
IJCNN 2025 · THU B, CCF C4th author · 3rd student author

Research FocusLLM Evaluation & Benchmarking

TechSupportEval: An Automated Evaluation Framework for Technical Support Question Answering

Bohan Chen, Yongqian Sun, Yuhe Liu, Longlong Xu, Zhe Xie, Changhua Pei, Jing Han, Fan Ni, Xuhui Cai, Ce Yang, Dan Pei

Automates technical-support QA evaluation with key-term matching, step-order verification, and completeness checks, outperforming the previous state of the art by 7.6% AUC.

Adversarial textured 3D meshes framework figure.
CVPR 2023 · CCF A · Highlight · 79+ citations3rd author

Research FocusRobust AI Systems

Towards Effective Adversarial Textured 3D Meshes on Physical Face Recognition

Xiao Yang, Chang Liu, Longlong Xu, Yikai Wang, Yinpeng Dong, Ning Chen, Hang Su, Jun Zhu

Develops adversarial textured 3D meshes that attack open-source face-recognition models across architectures and bypass real-world defense systems, including phone unlock and smart-door-lock systems.

Overall simulation framework for developing physical attacks figure.
IJCV 2024 · CCF A2nd author

Research FocusRobust AI Systems

Face3DAdv: Exploiting Robust Adversarial 3D Patches on Physical Face Recognition

Xiao Yang, Longlong Xu, Tianyu Pang, Yinpeng Dong, Yikai Wang, Hang Su, Jun Zhu

Optimizes face patches in the 3D StyleGAN latent space to achieve robust physical adversarial attacks against face-recognition systems.

Experience

Industry research experience across LLM post-training, multimodal reasoning, AIOps, numerical forecasting, and fault diagnosis.

01 LLM-Related Internships

ByteDance

Sep. 2025 - Aug. 2026

Research Intern. Built multimodal LLM post-training and evaluation pipelines for time-series reasoning, covering controllable data synthesis, multi-turn tool-use trajectories, SFT, Agentic RL, trajectory evaluation, RLVR, and efficient serving. Extended LLaMA-Factory, VeRL, TRL GRPOTrainer, and vLLM for native time-series inputs; work spans ChatTS, FoundRoot, IntentAD, and G2AD.

ZTE

Jan. 2024 - Aug. 2024

Research Intern. Built an end-to-end AIOps workflow for domain corpus and QA synthesis, LLM post-training, and automatic evaluation, improving domain QA accuracy by more than 80% over the chat base model. Co-authored OpsEval (FSE 2025) and TechSupportEval (IJCNN 2025); the public leaderboard has exceeded 2 million visits.

02 Other Internships

ByteDance

Jan. 2025 - Aug. 2025

Research Intern. Studied robustness and inference efficiency for time-series foundation models, addressing recent-observation shortcuts and long-context limitations. Produced two first-author papers, TameR and SPRINT (ICML 2026).

Huawei

Sep. 2024 - Dec. 2025

Research Intern. Studied multisource spatio-temporal foundation models for optical transport networks and automatic evaluation-data generation from hierarchical documents; contributed to Eagle (FSE 2026, CCF A) and supported engineering deployment.

eBay

Aug. 2023 - Dec. 2023

Research Intern. Developed root-cause localization for complex internal systems by modeling causal relationships among metrics in latent space and identifying root causes through intervention. The algorithm was engineered for practical use and produced MicroScope (KDD 2024, CCF A).

RealAI

Aug. 2022 - Jul. 2023

Research Intern. Studied physical-world 3D adversarial attacks for face-recognition model security, including adversarial textured meshes and 3D latent-space patch optimization. Co-authored AT3D (CVPR 2023, CCF A, Highlight) and Face3DAdv (IJCV 2024, CCF A).

Honors

Selected awards and recognitions.