●   Ph.D. candidate · University of Florida

Siqi Dai.

Building capable AI.
Making it trustworthy.

I work at the intersection of AI agents, multimodal learning, and privacy & security. My research and engineering span agent systems, vision-language model fine-tuning, and the evaluation of AI risks.

At ByteDance, I built and launched a self-improving agent system supporting 100+ workflows and 40+ users.

01 / Experience

From research to real systems.

ByteDance

E-commerce Group · San Jose, CA

May – Aug 2026

Machine Learning Engineer Intern

Built and launched a self-improving agent for e-commerce user-risk detection, integrating MCP tools with workflow execution traces for automated diagnosis and optimization.

100+workflows supported
94.3%recall, up from 63.1%
87%reduction in false-positive rate
  • Developed a CLI that gave the agent access to workflows and execution traces; scaled adoption to 40+ internal users.
  • Combined schema and tool validation, protected-node constraints, and LLM-as-a-Judge evaluation to validate candidate changes.
  • Reduced execution failures from 80% to near zero on targeted high-failure workflows using retries, fallbacks, and downstream validation.
  • Integrated the agent into a human-in-the-loop initiative ranked 5th among 40 projects in a ByteDance GNE department competition.

University of Florida

Gainesville, FL

Aug 2022 – Present

Research Assistant

Develop Python and PyTorch systems for VLM fine-tuning, multimodal privacy evaluation, adversarial robustness, federated learning security, and ML-based vulnerability analysis.

Build reproducible pipelines for model training, fine-tuning, and security evaluation, with peer-reviewed work in AI and security.

02 / Selected projects

Capabilities. Risks. Defenses.

A selection of systems and studies across conversational AI, vision-language models, and emerging interfaces.

Conversational AI · Privacy

PrivacyProbe

Automated privacy red teaming across four service personas. Structured conversational memory tracks inference evidence and guides adaptive questioning.

75.3% macro-averaged inference accuracy in the life-coach simulation across 30 synthetic profiles, 19 attributes, and 15-turn conversations.

Privacy red teamingStructured memory

Multimodal learning · Engineering

VLMs for PCB analysis

Fine-tuned Qwen2.5-VL with LoRA for EMI analysis using a multi-image, multi-turn dataset of 38 PCB designs and 254 images.

Combined text retrieval and device metadata in a three-stage reasoning pipeline, alongside a Gemini-based report generation pipeline.

Qwen2.5-VL / LoRARAG

AI security · Robustness

VLM jailbreak defense

A quantization-based defense against adversarial VLM jailbreaks, with a ViT–ResNet meta-learning detector for adaptation to emerging adversarial inputs.

95% defense success rate in the evaluated setting; studied cross-modal gradient alignment and transferability across open-source VLMs.

Adversarial robustnessMeta-learning

Spatial computing · Privacy

GAZEploit

Identified an Apple-confirmed privacy vulnerability in Apple Vision Pro: remote keystroke inference from avatar gaze observations.

Built an ML attack pipeline combining 3D gaze estimation, typing-session detection, and top-K prediction.

Read the paper ↗

Multimodal AI · Contextual privacy

ShareGuard

A four-stage image-sharing privacy review engine that separates sensitive-content extraction and audience-boundary inference from recipient matching.

Improved GPT-4.1 sharing-decision accuracy from 75.0% to 91.2%; benchmarked five VLMs on 160 cases.

Privacy evaluationVision-language models

Distributed ML · Security

Federated learning security

Evaluated targeted model poisoning that reduced target-class accuracy to 35% while limiting changes in non-target-class performance.

Used KL divergence to guide attack stealth and evaluated evasion of Krum and k-out-of-n selection defenses.

Read the paper ↗

03 / Selected publications

Research in print.

All publications on Scholar ↗
  1. PCIM 2026

    EMI Diagnosis for Power Converter PCB Layouts based on a Reasoning Aligned Vision Language Model ↗

    PCIM Conference · Conference program

  2. IEEE ICC 2025

    Guided by Noise: Vulnerable Poisoning Attack to Differentially Private Federated Learning ↗

    Siqi Dai, Yaodan Hu, Honggang Yu, Hanqiu Wang, and Shuo Wang

  3. ACM CCS 2024

    GAZEploit: Remote Keystroke Inference Attack by Gaze Estimation from Avatar Views in VR/MR Devices ↗

    Hanqiu Wang, Zihao Zhan, Haoqi Shan, Siqi Dai, Maximilian Panoff, and Shuo Wang

  4. ICRA Workshop 2024

    Zero-shot Safety Prediction for Autonomous Robots with Foundation World Models ↗

    Zhenjiang Mao, Siqi Dai, Yuang Geng, and Ivan Ruchkin

    Back to the Future: Robot Learning Going Probabilistic Workshop

04 / Education

Learning & research.

Aug 2022 – May 2027 (expected)

University of Florida

Ph.D. candidate
Electrical & Computer Engineering

Sep 2018 – Jun 2022

Southern University of
Science and Technology

Bachelor of Engineering

05 / Toolkit

Tools of the trade.

Languages & infrastructure
Python, SQL, Bash/Shell, MATLAB, AWS, Linux, Git
Model training
PyTorch, Hugging Face Transformers, LLaMA-Factory, PEFT, LoRA, multimodal SFT
Agents & evaluation
MCP, CLI development, structured memory, LLM-as-a-Judge, schema validation, human-in-the-loop
Retrieval & APIs
LangChain, FAISS, Sentence Transformers, text RAG, Gemini API
AI security
Privacy red teaming, adversarial robustness, jailbreak defense, model poisoning

Get in touch

Let's connect.

For conversations about AI agents, multimodal learning,
and trustworthy AI systems.

dais@ufl.edu ↗