W

Wencong Zhang

About

Detail

California, United States

Contact Wencong regarding: 
work
Full-time jobs
Flexible work
id_card
Internships
person_search
Finding candidates
connect_without_contact
Finding mentors
Finding co-founders
groups
Networking

Timeline


work
Job
school
Education

Résumé


Jobs verified_user 0% verified
  • Apple
    Senior Machine Learning Engineer
    Apple
    Feb 2023 - Current (3 years 8 months)
    Core ML Runtime & Generative AI for Ads and Developer Productivity (Transformer Optimization, On-Device/Cloud Inference, Ad Targeting, Agentic CI/CD) with Privacy-Preserving RAG Pipelines, Guardrails, and Scalable Model Deployment • Architected and optimized Core ML's execution stack to support on-device generative AI - including text generation, summarization, and multimodal tasks across iPhone, iPad, and Mac. • Implemented transformer-specific runtime features including fused attention, rotary embeddings, and optimized layer norms using C++, Metal, and Accelerate, reducing average token latency by up to 40%. • Extended Core ML converter to support PyTorch- and ONNX-exported models, adapting dynamic attention patterns and position encod
  • Cresta
    Staff Software Engineer
    Cresta
    Dec 2021 - Jan 2023 (1 year 2 months)
    Real-Time Audio Intelligence Platform (Go, gRPC, ASR, GKE, Pub/Sub) with Sub-Second Agent Assist, Scalable Ingestion, and End to-End ML + Ops Integration • Architected and led development of Gowalter, a real- time audio intelligence platform built in Go (gRPC, Protobuf), handling audio ingestion, streaming ASR, utterance segmentation, and sub-500ms coaching signal delivery. • Engineered Go-based backend to sustain 10K+ concurrent audio streams by optimizing gRPC connection pools, fine- tuning GKE autoscaling, and configuring L7 load balancers for efficient traffic distribution. • Deployed Gowalter on GCP, leveraging Kubernetes (GKE), Google Pub/Sub, and autoscaling node pools to handle thousands of concurrent calls while maintaining stri
  • Apple
    Senior Software Engineer
    Apple
    May 2021 - Dec 2021 (8 months)
    Core ML Model Conversion Validation and Debugging Tooling (Precision Drift, Op Compatibility, Float16 Optimization) • Developed internal model validation pipelines to verify accuracy preservation during conversion from PyTorch/TensorFlow to Core ML, with early detection of silent mismatches in tensor shapes, quantization ranges, and op behavior. • Designed tools to trace conversion-time transformations across layers - helping internal teams diagnose discrepancies between source and Core ML inference outputs with bitwise precision. • Collaborated with compiler and Core ML runtime teams to validate op compatibility and optimize subgraphs prone to degradation during mixed- precision conversion. • Defined patterns for safe deployment of Cor
  • Google
    Senior Software Engineer / Tech Lead
    Google
    Nov 2015 - Mar 2021 (5 years 5 months)
    AutoML Vision and Image Search Infrastructure (NAS, GPU Training, Multi- Tenant Pipelines, Relevance Features) for Scalable ML and Product Ranking Systems • Led infrastructure design for AutoML Vision, allowing users to train high-quality image classifiers without writing TensorFlow code - driving wide enterprise adoption of GCP AI. • Architected distributed pipelines for neural architecture search (NAS), dataset ingestion, multi-GPU training, and model export - supporting complex workflows with strict SLA guarantees. • Deployed autoscaling GPU-backed training services on Borg, using preemption-tolerant scheduling and bin-packing to maximize resource efficiency across multi-tenant clusters. • Built multi-tenant orchestration queues with
  • Amazon
    Software Development Engineer
    Amazon
    Sep 2014 - Oct 2015 (1 year 2 months)
    Backend Services for Order & Inventory Systems (Java, Spring, DynamoDB, SQS) with Resilient Fulfillment and High-Traffic Reliability at Amazon Scale • Developed and maintained core backend services in Java/Spring, powering inventory management, order routing, and customer tracking for Amazon's high-volume consumables business. • Designed RESTful APIs and data models to support multi-system workflows across fulfillment, availability, and personalized reordering. • Developed a fulfillment exception tracking system using DynamoDB, SQS, and Lambda, enabling the team to process and re queue failed order events without manual intervention - improving recovery time and order flow reliability during peak traffic. • Owned on-call rotations for l
Education verified_user 0% verified
  • Texas A&M University
    Master of Science
    Texas A&M University
    Sep 2012 - May 2014 (1 year 9 months)
  • Southeast University
    Bachelor of Science
    Southeast University
    Sep 2008 - May 2012 (3 years 9 months)