arrow_back 検索に戻る
シェア:
placeSydney, Australia home_work出社 labelAI & Applications public集約求人 · DE

event2026年9月13日に公開 · verified2026年9月14日時点で募集中であることを確認済みです

この企業はあなたの会社ですか?

求人について

Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally. Firmus AI Cloud Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers. It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale. Why Firmus? As an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience, and exposure across the AI industry and be part of shaping what this industry looks like for decades to come. We are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible, and there's minimum bureaucracy between you and the work. Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap.

Work alongside founders and experts in AI infrastructure, energy systems and next-generation compute. What we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in rather than drawing from them. Considering applying? You don't need a perfect background to join our team. If you're driven and curious, there's a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience. ROLE SUMMARY The Senior AI Engineer (Inferencing) will build and improve the AI & Applications team’s inference capability, making models available as reliable, secure, scalable, and high-performance endpoints for internal products, external customers, and future Inference-as-a-service offerings.

The role

will establish the engineering foundation for self-hosted model serving in the organization’s AI-factory environment. This includes model onboarding, deployment, endpoint provisioning, runtime selection, performance benchmarking and optimization, observability, capacity management, security, and operational lifecycle management. The objective is to provide users with predictable and efficient access to models while maintaining control over performance, cost, data handling, deployment configuration, and infrastructure utilization.

The role

is a key contributor to the Model-to-Grid product and agentic applications roadmap. It will convert model and runtime characteristics into benchmarked, repeatable inference recipes and endpoint profiles that can inform workload scheduling, topology-aware placement, capacity planning, performance recommendations, and operational decision-making. It will also provide the governed and fit-for-purpose model endpoints needed by agentic systems for reasoning, retrieval, tool use, diagnosis, recommendation, and controlled automation. KEY RESPONSIBILITIES Build, operate, and continuously improve self-hosted AI inference services for internal applications, customer-facing products, and future Inference-as-a-service offerings. Define and implement standard model-onboarding workflows covering model intake, compatibility validation, packaging, runtime selection, optimization, deployment, endpoint registration, testing, release, and lifecycle management. Provision and manage secure, scalable inference endpoints for common AI application patterns, including interactive generation, RAG, embeddings, reranking, batch processing, multimodal use cases, tool calling, and agentic workflows. Develop reusable deployment templates, APIs, SDKs, configuration standards, and self-service workflows for users to request, configure, access, monitor, update, and retire model endpoints. Work with leading inference frameworks and toolkits, such as TensorRT-LLM, TensorRT, SGLang, vLLM, Triton Inference Server, NVIDIA Dynamo, NVIDIA NIM, CUDA, cuDNN, NCCL, and related serving, profiling, and observability tools. Optimize model-serving performance using appropriate techniques, including quantization, compilation, batching, continuous batching, request routing, KV-cache management, prefix caching, speculative decoding, load balancing, model routing, memory optimization, and distributed parallelism. Build and validate reusable inference recipes that specify compatible model versions, framework and runtime versions, precision formats, GPU configurations, topology requirements, scaling approaches, scheduler profiles, benchmark results, and expected performance envelopes. Use quantization and optimization approaches such as NVFP4, FP8, INT8, TensorRT compilation, kernel optimization, efficient attention mechanisms, and memory-management techniques while maintaining agreed model-quality targets. Design distributed inference configurations for large models, including tensor, pipeline, expert, context, and data parallelism where appropriate. Work with the Kubernetes and proprietary scheduler team to define endpoint resource profiles, placement requirements, topology preferences, priority classes, quota models, autoscaling rules, capacity reservations, and workload-management policies. Contribute infe

このまま無料で続きを読む

無料アカウントを作成すると、求人の全文を見て応募できます。

  • badge企業からポートフォリオが見える
  • notifications新着求人をメールでお知らせ
  • favorite常に無料、追加料金なし