Xiaotong Wu
Xiaotong Wu outdoors at a scenic overlook

Xiaotong Wu

Master of Science in Engineering, San José State University

3D computer vision · Multimodal perception · Autonomous systems

I work on 3D computer vision and LiDAR–camera fusion, with experience in pillar-level fusion and BEV Transformer architectures for object detection. My research interests include robust and efficient multimodal perception, preserving temporal information for motion estimation, and adapting fusion to sensor reliability in autonomous systems.

I am a Research Assistant at Illinois iRisk Lab, University of Illinois Urbana-Champaign, advised by Zhiyu (Frank) Quan. I earned my Master of Science in Engineering at San José State University, where I studied interdisciplinary engineering with Ahmed Hambaba.

Research

LayerGate-BEV

Jul. 2026–Present

Illinois iRisk Lab · University of Illinois Urbana-Champaign

I proposed a scene-conditioned fusion extension that predicts camera–LiDAR weights for four BEVFormer encoder layers from pooled modality features. The proposal and architecture design are complete. I designed baseline-preserving gate initialization and a nuScenes evaluation plan covering detection, gate distributions, modality reliance, and robustness under LiDAR thinning.

BEVFormerFusion & FusionPointPillars

2024–2026

Master’s research · San José State University · Advisor: Ahmed Hambaba

I designed and implemented encoder-only and encoder–decoder LiDAR fusion for BEVFormer using deformable cross-attention and feature concatenation/projection, and conducted training, validation, and testing. On the full official nuScenes validation split, with matched baseline and fusion training settings on personal hardware, encoder–decoder fusion improved mAP from 0.201 to 0.251 and NDS from 0.219 to 0.255; mATE decreased from 0.949 to 0.870.

I developed LiDAR-branch fusion for FusionPointPillars and contributed to Adaptive Pillar Fusion architecture, implementation, and evaluation. On KITTI Moderate 3D AP40, GMF improved Pedestrian and Cyclist AP by 3.02 and 7.12 points, respectively; APF improved Cyclist AP by 7.38 points over the corresponding baselines.

End-to-end BEVFormerFusion inference on 200 samples, including data loading and preprocessing, reached 2.32 FPS for encoder–decoder fusion and 3.71 FPS for decoder-only fusion, compared with 3.86 FPS for the camera-only baseline, on a personal RTX 3060. I also authored model-description and experimental-results sections of manuscripts with my coauthors.

Publications & manuscripts

  • S. H. Suh, P. M. Girithimmappa, X. Wu, L. Shen, and A. Hambaba.

    Pillar-level multimodal fusion for efficient LiDAR–camera 3D object detection.

    WI-IAT 2026 · Accepted; camera-ready submitted

Submitted manuscripts

  • S. H. Suh, P. M. Girithimmappa, X. Wu, L. Shen, and A. Hambaba.

    FusionPointPillars: Efficient pillar-wise LiDAR–camera fusion for real-time 3D object detection in autonomous vehicles.

    IEEE Transactions on Intelligent Vehicles · Revision in progress

  • S. H. Suh, P. M. Girithimmappa, X. Wu, L. Shen, and A. Hambaba.

    Adaptive pillar fusion for efficient LiDAR–camera 3D object detection.

    IEEE ITSC 2026 · Revised manuscript submitted

Selected projects

Research & academic projects

Multimodal Restaurant Recommendation

2025

CMPE 256 team research project · San José State University

I loaded raw data, extracted BERT text and ResNet-18 image features, constructed community graphs using Louvain and FAISS, and trained a multimodal Graph Attention Network recommender. The model achieved validation RMSE of 0.1493 on ratings normalized to [0, 1], with early stopping for checkpoint selection.

Graph Link Prediction on ogbl-collab

2024

CS276 team project · San José State University

I co-designed a GAT encoder and pairwise MLP link predictor for a large coauthorship graph and implemented training and evaluation in PyTorch Geometric. Using OGB predefined splits and evaluator, binary cross-entropy, and negative sampling, the model achieved 46.44% test Hits@50 and 56.60% validation Hits@50.

Road Object Detection on Huawei Cloud

2020

Undergraduate team project · Changchun University of Technology

I trained a TensorFlow model for road-scene object detection on Huawei Cloud, targeting vehicles, pedestrians, road signs, and surrounding buildings.

Software projects

Muses

2026

Personal project · Native music applications

I architected a modular macOS music application in Swift and SwiftUI, with SwiftData persistence, AVFoundation playback, and isolated web-session access through a helper with versioned IPC. Features include persistent queues, media caching, lyrics, and native controls, supported by automated tests and signing/notarization workflows. The macOS release is available; iOS is in beta and Windows is in development with C#/.NET and Avalonia.

OpenJam

2025

Personal project · Real-time collaborative whiteboard

I built a Go/Gin WebSocket backend with room-scoped broadcasting and Redis pub/sub for inter-server messaging, paired with a React/TypeScript client using optimistic updates. OpenJam integrates PostgreSQL persistence, MinIO storage, session authentication, and autosave, with the frontend embedded in a Go binary and Docker deployment workflows.

Education

San José State University

Aug. 2024–May 2026

Master of Science in Engineering

Interdisciplinary Engineering · Advisor: Ahmed Hambaba

Master’s project: From Pillars to Transformers: Two Fusion Paradigms Based on PointPillars for 3D Object Detection (May 2026).

Changchun University of Technology

Sep. 2018–Jun. 2022

B.S., Computer Science and Engineering

Professional experience

Software Engineer · Internal Tools & IT Systems

Aug. 2025–May 2026

San José State University · University Housing Services

I built a Next.js/React/TypeScript and MongoDB equipment checkout system with REST APIs, indexed search, JWT access, and PDF receipt archival. I containerized the system, documented backup and migration workflows, supported 300+ end-user devices and 100+ Apple devices, and maintained AWS-hosted StarRez exports.

Software Engineer · Quantitative Systems

Mar. 2024–Aug. 2024

Founder Securities Co., Ltd.

I developed Python/Pandas/NumPy components for signal generation, strategy execution, and portfolio/performance analytics. I implemented Backtrader backtesting and market-data ETL for repeatable strategy evaluation and monitored trading workflows with risk controls.

Technical skills

Deep learning & vision: PyTorch, TensorFlow, MMCV, MMDetection3D, CNNs, Transformers, BERT.

Graph learning & data: GAT, Louvain, FAISS, scikit-learn, NumPy, Pandas.

Programming: Python, Go, Swift, TypeScript/JavaScript, C#, SQL.

Systems & tools: Docker, Git, GitHub Actions, AWS, PostgreSQL, MongoDB, Redis, SQLite.