|
Advantest is seeking a highly experienced Senior Principal AI Systems Engineer to architect and deliver advanced production AI capabilities. This role is for a hands-on technical leader with deep expertise in agentic AI, open-source model development, model fine-tuning, evaluation systems, and reliable AI infrastructure. The ideal candidate combines advanced AI knowledge with strong software-engineering discipline and has experience building systems that reason across multiple steps, use tools, evaluate results, recover from failures, and improve from feedback. Key Responsibilities Agentic AI Systems * Architect production-grade systems for tool-using and multi-step AI agents. * Design orchestration, planning, memory, session management, retries, timeouts, observability, and failure recovery. * Establish structured and typed interfaces between models, tools, data services, and applications. * Build controls that prevent agents from bypassing required validation and approval stages. * Develop reusable agent frameworks that support multiple products and deployment environments. Model Development and Fine-Tuning * Fine-tune open-source language models and specialized models for complex technical applications. * Apply LoRA, QLoRA, full fine-tuning, distillation, preference optimization, and related post-training methods. * Develop efficient strategies for adapting foundation models to new applications and datasets. * Evaluate tradeoffs among model quality, inference performance, deployment cost, security, and maintainability. * Build optimized models and inference profiles for constrained deployment environments. Evaluation and Quality * Define measurable standards for model accuracy, reliability, safety, and production readiness. * Build offline evaluation datasets, automated regression suites, judge systems, and promotion gates. * Develop methods for measuring confidence, consistency, tool-use accuracy, and action quality. * Establish processes for model comparison, controlled release, rollback, and continuous improvement. * Ensure model outputs remain grounded in available evidence and approved data sources. Training Data and Learning Pipelines * Design reproducible pipelines for training-data generation, cleaning, labeling, versioning, and validation. * Develop synthetic-data and preference-data strategies where appropriate. * Implement controls for data leakage, contamination, duplication, provenance, and customer isolation. * Convert expert feedback and observed outcomes into high-quality training and evaluation datasets. * Maintain traceability between datasets, experiments, model versions, and production results. AI Platform and MLOps * Build repeatable training, evaluation, and deployment workflows for multi-GPU infrastructure. * Establish model lifecycle practices, including model cards, release criteria, monitoring, and rollback. * Support secure, private, air-gapped, and customer-controlled deployment environments. * Develop observability for model behavior, tool execution, latency, cost, and failure conditions. * Partner with platform and CI/CD teams to make AI workflows repeatable, testable, and auditable. Technical Leadership * Set technical direction for a small, highly skilled AI engineering team. * Review architectures, models, training methods, and production implementation decisions. * Mentor engineers and establish durable AI engineering practices. * Work with domain experts to translate complex technical requirements into reliable AI capabilities. * Communicate technical risks, tradeoffs, progress, and recommendations to engineering and executive stakeholders.
|