Xiaotong Wu
Master of Science in Engineering, San José State University
3D computer vision · Multimodal perception · Autonomous systems
I work on 3D computer vision and LiDAR–camera fusion, with experience in pillar-level fusion and BEV Transformer architectures for object detection. My research interests include robust and efficient multimodal perception, preserving temporal information for motion estimation, and adapting fusion to sensor reliability in autonomous systems.
I am a Research Assistant at Illinois iRisk Lab, University of Illinois Urbana-Champaign, advised by Zhiyu (Frank) Quan. I earned my Master of Science in Engineering at San José State University, where I studied interdisciplinary engineering with Ahmed Hambaba.
Research
LayerGate-BEV
Jul. 2026–Present
Illinois iRisk Lab · University of Illinois Urbana-Champaign
I proposed a scene-conditioned fusion extension that predicts camera–LiDAR weights for four BEVFormer encoder layers from pooled modality features. The proposal and architecture design are complete. I designed baseline-preserving gate initialization and a nuScenes evaluation plan covering detection, gate distributions, modality reliance, and robustness under LiDAR thinning.
BEVFormerFusion & FusionPointPillars
2024–2026
Master’s research · San José State University · Advisor: Ahmed Hambaba
I designed and implemented encoder-only and encoder–decoder LiDAR fusion for BEVFormer using deformable cross-attention and feature concatenation/projection, and conducted training, validation, and testing. On the full official nuScenes validation split, with matched baseline and fusion training settings on personal hardware, encoder–decoder fusion improved mAP from 0.201 to 0.251 and NDS from 0.219 to 0.255; mATE decreased from 0.949 to 0.870.
I developed LiDAR-branch fusion for FusionPointPillars and contributed to Adaptive Pillar Fusion architecture, implementation, and evaluation. On KITTI Moderate 3D AP40, GMF improved Pedestrian and Cyclist AP by 3.02 and 7.12 points, respectively; APF improved Cyclist AP by 7.38 points over the corresponding baselines.
End-to-end BEVFormerFusion inference on 200 samples, including data loading and preprocessing, reached 2.32 FPS for encoder–decoder fusion and 3.71 FPS for decoder-only fusion, compared with 3.86 FPS for the camera-only baseline, on a personal RTX 3060. I also authored model-description and experimental-results sections of manuscripts with my coauthors.
BEVFormerFusion codeFusionPointPillars code
Publications & manuscripts
-
S. H. Suh, P. M. Girithimmappa, X. Wu, L. Shen, and A. Hambaba.
Pillar-level multimodal fusion for efficient LiDAR–camera 3D object detection.
WI-IAT 2026 · Accepted; camera-ready submitted
Submitted manuscripts
-
S. H. Suh, P. M. Girithimmappa, X. Wu, L. Shen, and A. Hambaba.
FusionPointPillars: Efficient pillar-wise LiDAR–camera fusion for real-time 3D object detection in autonomous vehicles.
IEEE Transactions on Intelligent Vehicles · Revision in progress
-
S. H. Suh, P. M. Girithimmappa, X. Wu, L. Shen, and A. Hambaba.
Adaptive pillar fusion for efficient LiDAR–camera 3D object detection.
IEEE ITSC 2026 · Revised manuscript submitted
Selected projects
Research & academic projects
Multimodal Restaurant Recommendation
2025
CMPE 256 team research project · San José State University
I loaded raw data, extracted BERT text and ResNet-18 image features, constructed community graphs using Louvain and FAISS, and trained a multimodal Graph Attention Network recommender. The model achieved validation RMSE of 0.1493 on ratings normalized to [0, 1], with early stopping for checkpoint selection.
Restaurant recommendation code
Graph Link Prediction on ogbl-collab
2024
CS276 team project · San José State University
I co-designed a GAT encoder and pairwise MLP link predictor for a large coauthorship graph and implemented training and evaluation in PyTorch Geometric. Using OGB predefined splits and evaluator, binary cross-entropy, and negative sampling, the model achieved 46.44% test Hits@50 and 56.60% validation Hits@50.
Road Object Detection on Huawei Cloud
2020
Undergraduate team project · Changchun University of Technology
I trained a TensorFlow model for road-scene object detection on Huawei Cloud, targeting vehicles, pedestrians, road signs, and surrounding buildings.
Software projects
Personal project · Native music applications
I architected a modular macOS music application in Swift and SwiftUI, with SwiftData persistence, AVFoundation playback, and isolated web-session access through a helper with versioned IPC. Features include persistent queues, media caching, lyrics, and native controls, supported by automated tests and signing/notarization workflows. The macOS release is available; iOS is in beta and Windows is in development with C#/.NET and Avalonia.
Muses code
Personal project · Real-time collaborative whiteboard
I built a Go/Gin WebSocket backend with room-scoped broadcasting and Redis pub/sub for inter-server messaging, paired with a React/TypeScript client using optimistic updates. OpenJam integrates PostgreSQL persistence, MinIO storage, session authentication, and autosave, with the frontend embedded in a Go binary and Docker deployment workflows.
OpenJam code
Education
San José State University
Aug. 2024–May 2026
Master of Science in Engineering
Interdisciplinary Engineering · Advisor: Ahmed Hambaba
Master’s project: From Pillars to Transformers: Two Fusion Paradigms Based on PointPillars for 3D Object Detection (May 2026).
Changchun University of Technology
Sep. 2018–Jun. 2022
B.S., Computer Science and Engineering
Professional experience
Software Engineer · Internal Tools & IT Systems
Aug. 2025–May 2026
San José State University · University Housing Services
I built a Next.js/React/TypeScript and MongoDB equipment checkout system with REST APIs, indexed search, JWT access, and PDF receipt archival. I containerized the system, documented backup and migration workflows, supported 300+ end-user devices and 100+ Apple devices, and maintained AWS-hosted StarRez exports.
Software Engineer · Quantitative Systems
Mar. 2024–Aug. 2024
Founder Securities Co., Ltd.
I developed Python/Pandas/NumPy components for signal generation, strategy execution, and portfolio/performance analytics. I implemented Backtrader backtesting and market-data ETL for repeatable strategy evaluation and monitored trading workflows with risk controls.
Technical skills
Deep learning & vision: PyTorch, TensorFlow, MMCV, MMDetection3D, CNNs, Transformers, BERT.
Graph learning & data: GAT, Louvain, FAISS, scikit-learn, NumPy, Pandas.
Programming: Python, Go, Swift, TypeScript/JavaScript, C#, SQL.
Systems & tools: Docker, Git, GitHub Actions, AWS, PostgreSQL, MongoDB, Redis, SQLite.