Ke Xu

Ke Xu 许可

Algorithm Engineer II, Huawei 2012 Laboratory

About

I am Ke Xu, an Algorithm Engineer II at Huawei 2012 Laboratory. I started with RLHF training infrastructure (GRPO and DAPO on Ascend NPU clusters), and my focus has since moved to agent systems: multi-agent scheduling for scientific discovery and, most recently, recursive self-improvement (RSI).

I completed my M.S. in Computer Science at the Computer Science and Engineering Department, University of California San Diego (UCSD), where my concentration was Machine Learning.

Prior to that, I graduated from a dual-degree program jointly held by Zhejiang University and the University of Illinois at Urbana-Champaign, with major in Electrical Engineering and minor in Computer Science, where I was a member of Engineering National Graduate Institutional Name Exchange (ENGINE), a participant of Research Experiences for Undergraduate (REU) supported by National Science Foundation (NSF) 1947135, advised by Prof. Hanghang Tong.

Previously, I was an Algorithm Engineer Intern at Ant Group, where I worked on the application of Large Language Models (LLM) in security and risk management domains such as Anti-Money Laundering (AML).

Research Interests: Reinforcement learning and agent systems, especially recursive self-improvement (RSI).

Latest News

Work Experience

Algorithm Engineer II May 2025 – Present
Huawei 2012 Laboratory
Large-Scale RLHF Training
  • Implemented core modules of the GRPO and DAPO algorithms for large-scale RLHF training on Ascend NPU clusters.
  • Responsible for GRPO training and tuning of large models, including dynamic sequence packing for long chain-of-thought samples and stability validation at cluster scale.
Agent System for Scientific Discovery
  • Participated in the architecture design and core development of a scientific-discovery agent framework; built its pluggable DAG-based multi-agent scheduler with configurable node types and dependencies.
  • Building the RSI loop into the framework: agents evolve their own skill libraries on task datasets through rollout, failure-driven skill induction, and top-K frontier selection, under strict train/validation/test discipline; every evolved library is versioned and auditable.
  • Exploring open-ended evolution, with island populations and MAP-Elites quality-diversity archives alongside tree search; next, self-generated evaluation tasks so agents can drive their own curriculum.
Algorithm Engineer Intern Jun 2023 – Sep 2023
Ant Group
  • Built a financial agent prototype on AntGLM and LangChain: ReAct-style reasoning, hybrid BM25/embedding retrieval, and multi-turn tool calling.
  • Fine-tuned AntGLM with LoRA for anti-money-laundering (AML) anomaly detection.
AI Engineer Intern Nov 2022 – Feb 2023
State Street
  • Designed the reconciliation algorithms for an AI-based automatic reconciliation system.

Research Experience

Active Learning on Knowledge Graphs Question Answering Summer 2022
Advisor: Prof. Hanghang Tong, University of Illinois at Urbana-Champaign

Designed an active-learning strategy for KGQA that fuses graph centrality, information density, and node uncertainty.

Exploration of Reflective ASMs for Security Apr 2021 – Apr 2022
Advisor: Prof. Klaus-Dieter Schewe, Zhejiang University

Investigated reflective Abstract State Machines (ASMs) for security.

AI-Assisted Psychological Counseling Chatbot Summer 2021
Advisor: Dr. Zhenzhong Lan, Westlake University

Built blacklist/whitelist safety classification and knowledge-graph-based intent and emotion recognition for counseling dialogue.

Education

Honors & Awards

Travel

Away from the terminal, I love exploring the world — 0 countries and counting. visited   planned   not yet

Full checklist — 0 of 0 countries visited