وصف الوظيفة
Fuse Energy هي شركة ناشئة رائدة في مجال الطاقة المتجددة في مقدمة التفكير، في مهمة لتسليم تيراوَاط من الطاقة المتجددة - بسرعة.
نحن نجمع بين التفكير من المبادئ الأولى والتكنولوجيا المتقدمة لبناء نظام طاقة radically أفضل.
جمعنا 210 مليون دولار من مستثمرين من المستوى الأول بما في ذلك Multicoin وBalderton وLakestar وAccel وCreandum وLowercarbon وRibbit وBox Group وملاك استراتيجيين مثل نيكو روسبرغ، المؤسس المشارك لـ Solana وشركاء الاستثمار وراء Meta وRevolut وSpotify وUber وغيرهم.
مع أن مراكز البيانات تصبح واحدة من أكبر وأكثر مصادر الطلب على الكهرباء نمواً بسرعة، فإن Fuse تتوسع في بنية تحتية للحوسبة عالية الأداء تقطع تقاطع الطاقة والذكاء الاصطناعي.
نحن نبني طبقة أداء GPU/CUDA وخدمة الاستدلال في نفس الوقت، من الصفر - ونحن نبحث عن مهندس تأسيسي ليكون مالك الأخيرة.
نحن نبحث عن مهندس استدلال ذكاء اصطناعي تأسيسي لتعريف وبناء كيف تقدم Fuse أحمال استدلال الذكاء الاصطناعي على نطاق واسع، مع الإبلاغ مباشرة إلى CTO.
حيث يملك موظفو هندسة CUDA وGPU مستوى النواة وأداء الأجهزة، تملك هذه الوظيفة الطبقة التي فوقها: كيف يتم تقديم النماذج فعلياً، وتكبيرها، وتوصيلها إلى أداء مُلتزم.
الفرصة: تشهد Fuse طلباً كبيراً على سعة مراكز البيانات في الأسواق التي نعمل فيها، أساساً للاستدلال.
قلة من الشركات في العالم يمكنها مزج تقديم طاقة حقيقية مع حوسبة حقيقية بالطريقة التي يمكن لـ Fuse، مما يجعل خدمة الاستدلال في قلب كيف نحول تلك الميزة إلى أفضل عرض في السوق.
هذا هو هذا المنصب.
المسؤوليات: تحديد استراتيجية وبنية خدمة الاستدلال في Fuse من المبادئ الأولى.
تصميم وبناء طبقة الخدمة: توجيه الطلبات، التجميع، الجدولة، والتوسع التلقائي لأعباء الاستدلال العالية الإنتاجية والدلالية من حيث الكمون.
امتلاك استراتيجية تحسين على مستوى النموذج للخدمة - تحديد أين وكيف يتم تطبيق التقريب الكمي، التقطير، فك التشفير المضارب، وتقنيات مماثلة لتحسين معدل النقل والتكلفة لكل رمز، بالشراكة مع مهندسي CUDA/GPU.
إجراء قرارات هندسية أساسية حول أطر وخدمات الاستدلال والتنسيق (مثلاً vLLM وTensorRT-LLM وSGLang وTriton Inference Server أو ما يعادلها).
تحويل الالتزامات حول معدل النقل والكمون والتوافر إلى مواصفات تقنية وخطط قدرة خدمة.
العمل كمالك تقني مباشر لأداء الاستدلال وموثوقيته.
العمل عن كثب مع فرق هندسة CUDA وGPU لضمان تكامل النوى المخصصة وأعمال أداء الأجهزة في طبقة الخدمة بشكل نظيف.
وضع المعايير والأدوات والمرجعيات التي ستعمل عليها هذه الوظيفة مع نموها.
راتب تنافسي ومكافأة توقيع أسهم.
مخطط مكافأة نصف سنوي.
معدات تقنية متكاملة لتحسين احتياجاتك.
بدل إفطار وعشاء للموظفين العاملين في المكتب.
4+ سنوات من الخبرة في بناء أو تشغيل أنظمة خدمة استدلال على نطاق واسع، أو خبرة عملية/مشروع قوية مكافئة.
خبرة عميقة ويدوية في أطر استدلال وخوارزميات تحسينها (التجميع، إدارة KV-cache، التكميم، فك التشفير المضارب).
تفكير منظم قوي - القدرة على التفكير في المسار الكامل من الطلب الوارد إلى الرد المقدم عبر عقدة كبيرة.
القدرة على العمل مباشرة مع مهندسي GPU/CUDA لدمج أعمال الأداء منخفضة المستوى في نظام خدمة.
سجل من اتخاذ قرارات هندسية عالية المخاطر وتحمّل النتائج.
الراحة في العمل بدون دليل تشغيلي - هذه وظيفة تأسيسية تشكل وظيفة جديدة حول بنية تقويم لا تزال في مرحلة مبكرة، وليست الانضمام إلى واحدة قائمة.
من الجميل وجود تجربة مع Triton أو أطر استدلال/تدريب ML مخصصة.
خبرة مع التوسع التلقائي أو تخطيط السعة لحمل استدلال واسع النطاق.
التعرض لخدمة متعددة المستأجرين أو بنية تحتية مدفوعة بخدمات مستوى SLA.
خلفية في مشغل فائق أو مُختبر AI رائد أو نظام استدلال موزّع على نطاق واسع.
الإلمام بـ Kubernetes/Slurm من أجل تنظيم المجموعات.
اهتمام أو خبرة في أسواق الطاقة، أنظمة الشبكة، أو حوسبة مُستدامة.
Job description
Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast.
We're combining first-principles thinking with cutting-edge technology to build a radically better energy system.
We raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels like Nico Rosberg, the Co-Founder of Solana and GPs behind Meta, Revolut, Spotify, Uber and more.
As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI.
We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch - and we're looking for the founding engineer to own the latter.
We're looking for a Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO.
Where our CUDA and GPU engineering hires own kernel-level and hardware performance, this role owns the layer above it: how models actually get served, scaled, and delivered against committed performance targets.
The Opportunity Fuse is seeing significant demand for data centre capacity across the markets we operate in, primarily for inference.
Few companies in the world can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of how we turn that advantage into the best offering in the market.
That's this role.
Responsibilities Define Fuse's inference serving strategy and architecture from first principles.
Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads.
Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers.
Make the core software architecture calls on serving frameworks and orchestration (e.
g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents).
Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans.
Act as a direct technical owner of inference performance and reliability.
Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer.
Set the standards, tooling, and benchmarks this function will run on as it grows.
Competitive salary and an equity sign-on bonus.
Biannual bonus scheme.
Fully expensed tech to match your needs.
Breakfast and dinner allowance for office based employees.
4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience.
Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding).
Strong systems thinking - able to reason about the full path from incoming request to served response across a large cluster.
Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system.
A track record of making high-stakes architecture calls and owning the outcome.
Comfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one.
Nice to Have Experience with Triton or custom ML inference/training frameworks.
Experience with autoscaling or capacity planning for large-scale inference workloads.
Exposure to multi-tenant serving or SLA-driven infrastructure.
Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system.
Familiarity with Kubernetes/Slurm for cluster orchestration.
Interest or experience in energy markets, grid systems, or sustainability-focused compute.