Senior Software Engineer, Inference Engine (Platform Software)
Seoul, South KoreaCompilersKernel / LinuxPosted Aug 19, 2026via greenhouse
About FuriosaAI
FuriosaAI builds high-performance, high-efficiency AI compute for the Inference Era. Founded in 2017 by veteran semiconductor and AI algorithm engineers, Furiosa operates globally with offices in Korea and Silicon Valley, along with a compiler-focused R&D lab in Lisbon.
Our vision is to make AI computing sustainable, enabling access to powerful AI for everyone on Earth. We solve the AI hardware energy and operational cost crisis at the architectural level, rather than through brute force, building the world's first truly AI-native compute platform to unlock the full potential of artificial intelligence for every enterprise.
About the Role
Software Engineer (Inference Engine) is responsible for developing and optimizing a high-performance inference engine for Large Language Models (LLMs) and multimodal LLMs running on FuriosaAI NPUs.
In this role, you will proactively research and apply the state-of-the-art inference optimization techniques to our inference engine. You will work in close collaboration with the compiler and hardware teams to enhance the engine's performance to its full potential.
Key Responsibilities
• Design and implement FuriosaAI’s next-generation inference engine for large and multimodal language models—comparable in capability to frameworks such as vLLM and SGLang—optimized for throughput, latency, and memory efficiency.
• Design and implement advanced inference optimizations—such as speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling—in our production inference engine.
• Design and develop capabilities for distributed and scalable inference, including prefill–decode (PD) and encode–prefill–decode (EPD) disaggregation, disaggregated speculative decoding, and hierarchical and external KV-cache storage such as HiCache and Mooncake.
• Collaborate closely with the Compiler team to co-design and optimize execution for FuriosaAI NPUs, improving system-level throughput, latency, and memory utilization.
• Proactively research, evaluate, and integrate state-of-the-art inference optimization techniques and key features of LLM serving frameworks into our production inference engine.
Minimum Qualifications
• BS degree in Computer Science, Engineering, or a related field, with at least 3 years of relevant industry experience, or equivalent practical experience
• Proficiency in Rust or C++ programming skill
• Knowledge and passion of deep learning, LLM, and/or generative AI models
• Excellent problem-solving and data analysis skills.
• Strong communication and collaboration skills.
Preferred Qualifications
• Experience in building inference serving systems for large models, encompassing batching, scheduling, caching, and load balancing.
• A deep understanding of performance optimization systems.
• Proficiency in C++/CUDA or Triton kernel development
• Contributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM.
Why Join FuriosaAI
The defining bottleneck of the AI era is building the right hardware and software stack to run it at global scale. Furiosa is solving this challenge holistically from the ground up.
With our flagship chip, RNGD, in mass production today and our next-generation platform in development with Broadcom, we are proving that full-stack, tensor-native compute is the future of AI infrastructure. This is a pivotal moment to join our team, right as we accelerate our global expansion.
At Furiosa, you will:
Solve AI’s Most Urgent Challenge. Help build the high-performance, energy-efficient inference hardware and software required to fulfill the promise of advanced AI.
Pioneer Full-Stack Co-Design. Work with teams that are architecting solutions from silicon up through the compiler (featuring innovations like Tensor Contraction Language and Virtual ISA) and serving frameworks.
Ship Real-World Silicon, Software, and Solutions. Turn breakthrough technology into commercial deployment. RNGD is in mass production with TSMC and running live enterprise workloads for global leaders like LG AI Research and Samsung SDS.
Partner With the Industry's Best. Collaborate across an elite global ecosystem that includes TSMC, Broadcom, SK Hynix, and GUC.
Do Your Life’s Best Work. Join a brilliant, low-ego, mission-driven team in a high-trust environment that values autonomy, intellectual curiosity, and shared ambition.
Contact
• recruit@furiosa.ai
Source URL: https://furiosa.ai/careers?gh_jid=4005788201