{"handle":"xinnyi","display_name":"Xin Yi","headline":"AI Engineer at Amazon","bio":"AI Engineer at Amazon focused on LLM agent, AI/ML systems, and distributed systems.\nExperience across embedding retrieval, model serving, high-throughput backend infrastructure, LLM post-training, and GPU performance optimization.\nInterested in AI Infrastructure, ML Systems, and LLM training and inference.","location":"Bay Area","company":"Amazon","role":"Software Engineer","open_to":"AI Platform / ML training & Inference related work","languages":"English, Chinese","skills":["Java","Python","C++/C","TypeScript","Kafka","Spark","PyTorch","TensorFlow","CUDA","Docker","Kubernetes","AWS","Machine Learning","LLM","Agent","Post Training","Machine Learning System"],"positions":[{"title":"Software Engineer","company":"Amazon","location":"Seattle","start_year":2025,"start_month":6,"end_year":null,"end_month":null,"description":"• Led EU + APAC expansion (19 marketplaces) of Audience Targeting Agent, an automated keyword/audience/category recommendation system for advertiser’s campaigns; architected split-region EU deployment to resolve conflict between data residency, Bedrock AgentCore and Ads Service availability with secure cross-region invocation. • Built marketplace-aware cross-lingual keyword retrieval by deploying a multilingual embedding model and re-ingesting 8.5M localized keyword records through new OpenSearch neural pipelines across three regions. • Delivered end-to-end observability for the multi-stage agent serving path using OpenTelemetry and CloudWatch; reduced token usage by 40% via compressed few-shot prompts, and cut TTFT by 27% using Gemma3-4B for ranking. • Co-designed and implemented the migration of audience-permission lookup and segment synchronization from a legacy graph-based model to a DynamoDB two-table design, increasing supported throughput from 2 to 10,000 TPS (5,000×)."},{"title":"Software Engineer Intern","company":"Amazon","location":"Seattle","start_year":2024,"start_month":5,"end_year":2024,"end_month":11,"description":"• Built a distributed user-activity dual-stage ingestion pipeline: GraphQL→SNS→SQS→Lambda→transaction DB with early filtering and elastic concurrency, plus a DynamoDB Streams-triggered aggregator Lambda with hybrid batching; sustained ∼150ms average latency and zero observed errors at 2× peak production load. • Designed retry and failure isolation mechanisms; automated multi-region deployment via AWS CDK (TypeScript) CI/CD with CloudWatch monitoring."},{"title":"Machine Learning Framework Engineer","company":"Huawei","location":"Remote","start_year":2022,"start_month":7,"end_year":2022,"end_month":10,"description":"Dilation CPU Operators Development for MindSpore Open-Source Framework • Designed Python front-end APIs and C++ back-end logic for Dilation and two backpropagation operators, supporting 12 tensor precisions with relative error under 0.002% vs. NumPy. • Optimized performance by flattening 3D tensors to reduce branch divergence and adapting static-shape inputs to dynamic shape for flexible memory management."},{"title":"Machine Learning Engineer Intern","company":"ML4SCI","location":"Remote","start_year":2022,"start_month":6,"end_year":2022,"end_month":8,"description":"• Focused on the problem of data sparsity in image-based particle recognition methods; represented particle jets images as graphs and explored various channel combinations as node features; determined the most appropriate representation strategies. • Defined edges in a static or dynamic manner and computed connectivity in static graphs using KNN or radius neighbors. • Built various GNN model architectures for end-to-end tau particle recognition using current cutting-edge methods in graph deep learning research, including graph convolution, graph SAGE, graph attention, and dynamic edge convolution. • Analyzed model performance from multiple perspectives and provided possible explanations. Benchmarked model inference on GPUs and provided guidance for users through documentation and code. • Developed CLI tools in Python to automate training and inference, improving development efficiency by 40%."},{"title":"Research Assistant","company":"Huazhong University of Science and Technology","location":"Wuhan, China","start_year":2023,"start_month":1,"end_year":2023,"end_month":4,"description":"GPU Runtime Scheduling for GNN Training • Built a persistent CUDA task-pool scheduler for GNNAdvisor with warp-level workload balancing, occupancy-aware execution, and kernel fusion to reduce launch and global-memory overhead. • Benchmarked GCN/GIN GPUkernels across 15 graph datasets, validating against PyG and achieving up to 5.3% speedup over hardware scheduling."}],"education":[{"school":"Texas A&M University","degree":"Master of Computer Science","field_of_study":null,"start_year":2023,"end_year":2025,"description":"GPA 3.8/4.0"},{"school":"Huazhong University of Science & Technology","degree":"Bachelor of Engineering","field_of_study":"Computer Science","start_year":2019,"end_year":2023,"description":"cGPA 3.7/4.0, Major 3.95/4.0"}],"work_email":"yixin5040@gmail.com","phone":"979-739-9282","calendar_url":null,"website_url":null,"linkedin_url":"https://linkedin.com/in/xinyii/","x_url":null,"github_url":"https://github.com/Allyyi","locked":[],"contact_visibility":"open","unlock_price_cents":null,"call_price_cents":null,"accepts_hiring_messages":true,"work_verified":true,"work_verified_domain":"amazon.com","is_public":true,"owned":false,"created_at":"2026-09-30T06:10:25.412918Z","agents":[],"agent_count":0,"follower_total":0,"posts":[],"posts_total":0}