Interviews

Kismat Singh, Co-Founder and CEO of MachGen AI – Interview Series

mm
Add Unite.AI to your preferred sources on Google

Kismat Singh, Co-Founder and CEO of MachGen AI, is an AI infrastructure and engineering executive with more than two decades of experience building high-performance computing and machine learning systems. Before co-founding MachGen AI, Singh served as VP of Engineering for AI Frameworks at Intel and previously held senior engineering roles at NVIDIA, AMD, HP, and other technology companies. His work has included contributions to deep learning libraries and compilers, AI framework optimization, and improvements to hyperscale training resiliency. At Intel, he was closely involved with the company’s PyTorch initiatives and spoke publicly about making AI frameworks more accessible across different hardware platforms.

MachGen AI is an AI infrastructure company focused on making generative image and video models dramatically faster and more economical to run. Founded by Kismat Singh and Manoj Krishnan, the company is developing a high-performance inference and fine-tuning stack specifically designed for diffusion models rather than adapting infrastructure originally built for large language models. MachGen optimizes areas including attention computation, caching, GPU kernels, memory management, parallelism, scheduling, and deployment to reduce generation latency while maintaining model quality. Its platform provides APIs for text-to-image, image-to-image, text-to-video, image-to-video, and reference-to-video generation, with the underlying technology aimed at enabling low-latency applications such as interactive content creation, real-time avatars, gaming, and personalized advertising.

You’ve spent more than two decades working on video, GPU compute, and AI performance, including TensorRT at NVIDIA, AI software at Intel, GPU optimization at AMD, and earlier work on real-time video encoding. What did you see during those years that ultimately convinced you to leave large technology companies and co-found MachGen, and why did you believe diffusion inference was the infrastructure problem worth building a company around?

My cofounder Manoj and I played badminton at the same gym for nearly 10 years. For most of that time, we were trying to hire the other one. Eventually we decided it was easier to work with each other than to keep recruiting each other.

The reason we picked this problem is that we’ve both watched an inference stack mature from the inside. I helped build NVIDIA’s inference platform team and then led AI software at Intel. Manoj built the large-model training infrastructure behind PyTorch at Meta. We watched years of work on caching, batching, scheduling, and kernels turn early LLM research systems into production infrastructure serving trillions of tokens.

Diffusion is roughly where LLM inference was three years ago. Image and video models have advanced rapidly, but the systems serving them remain slow, expensive, and difficult to scale. Most existing infrastructure was designed for LLMs, with diffusion added later.

Three things that made the timing right: the models became good enough to build real products on, the problem needed exactly the skills we’d spent our careers developing, and the impact is large enough to build a generational company around. When a video that once took several minutes can be generated in seconds at a fraction of the cost, interactive generation, personalized advertising, real-time avatars, and entirely new creative products become possible.

MachGen argues that many of the techniques that transformed large language model inference don’t transfer cleanly to diffusion models. What is fundamentally different about serving image and video models, and where are the biggest performance bottlenecks that conventional AI infrastructure tends to miss?

LLMs and diffusion models have fundamentally different computational patterns. LLMs are autoregressive, with a compute-heavy prefill followed by a memory-bound, token-by-token decode. Techniques such as KV caching and prefill/decode disaggregation were designed around that structure.

Diffusion models are non-autoregressive and remain compute-bound throughout the denoising process. Video generation also involves large amounts of spatial and temporal data, creating different demands around attention, memory, communication, and parallelism.

As a result, many of the optimizations developed for LLMs don’t transfer cleanly. Conventional KV caching doesn’t solve the same problem, standard parallelism can introduce substantial communication overhead, and the available open-source kernels are much less mature. The scheduler, memory model, caching strategy, kernels, and parallelism all need to be designed around diffusion itself. That’s why we built MachGen with diffusion as a first-class citizen.

MachGen reports roughly 4–6x lower latency on some image models and around 6x improvements on several video models while running the original, undistilled models. What are the most important technical breakthroughs behind those gains, and how do you ensure performance improvements don’t come at the expense of output quality?

The gains come from several parts of the stack working together. Attention dominates the compute budget in many video models, so we focus on lower-precision recipes, sparse attention, and related techniques. We’ve also developed caching methods that exploit spatial and temporal redundancy, highly optimized kernels for diffusion architectures, and parallelism techniques that reduce communication overhead.

Together, those improvements allow MachGen to deliver HiDream in one second versus six seconds and Flux in 1.5 seconds versus 6.2 seconds. For video, Wan 2.2 runs in 16.5 seconds versus 98 seconds, while LTX 2.3 runs in 10.7 seconds versus 67 seconds. Inference costs are also 2–4x lower for image models and 2–3x lower for video.

These results use the original, undistilled models. We’re improving how the models run without sacrificing output quality. For production adoption, speed, cost, and quality all have to work together.

Trusted TV is an interesting example because reducing generation from minutes to seconds changes more than the infrastructure bill. Once video generation becomes fast enough to produce thousands of campaigns and personalized creative variations, what new business models or product experiences become possible that simply weren’t viable before?

TrustedTV helps businesses create broadcast-ready commercials with AI and place them across connected TV platforms like Prime Video, Roku, and DirecTV. Many of its customers are small businesses. Making a TV commercial the traditional way meant a film and editing crew, with costs starting in the thousands and typically running into the tens of thousands of dollars. Trusted TV takes that to zero. A small business owner can create a full commercial from a smartphone in about five minutes and see it before paying anything. More than 10,000 campaigns have been created on the platform.

Ian, Trusted TV’s CEO, saw MachGen’s latency benchmark on Artificial Analysis and didn’t believe it. He signed up to test it himself. His first render came back in about 25 seconds with no loss in quality, and when he ran concurrency tests, the improvements held up at scale. They moved their entire production workflow to MachGen in under 48 hours, and MachGen now powers all of their new video creation. In the latest deployment, we’re closer to 10 or 11 seconds on that model, and we’ll keep improving it.

The product effect is that advertisers stop making one or two general commercials. Their users rotate two or three new ads per week, test creative across different audience segments, and iterate rather than commit. What comes next is geography and audience-level personalization at a scale, which was previously reserved for companies with multi-million-dollar budgets.

A version I think about personally: Today, I get a sneaker ad shot on a street in Manhattan. I live in the Bay Area, so it means nothing to me. What I want is the same ad with Mission Peak in the background. That’s not a cheaper commercial; it’s a different kind of advertising, and it only becomes economically viable when generation is fast and cheap enough to make thousands of variants.

For companies embedding image or video generation directly into products, how should they think about the economics of inference? At what scale does optimizing GPU utilization become strategically important rather than simply paying an API provider for each generation?

The inflection point comes when inference begins affecting either the customer experience or the company’s margins.

A general-purpose API can make sense while a team is validating a product. Once generation becomes central to that product, teams need to look beyond the listed price per output. Latency, concurrency, quality, reliability, and GPU utilization all become part of the economics.

A generation may appear inexpensive, but it still creates a business problem if users have to wait several minutes and abandon the workflow. An architecture that requires GPU capacity to grow at the same rate as usage will also put increasing pressure on margins.

There’s no single request-volume threshold because the compute required for a short 540p clip is very different from a longer 1080p video. Optimization becomes strategic when growing usage causes costs to rise at nearly the same rate, peak demand creates queues, or latency begins limiting the product experience.

MachGen works across several layers, including custom GPU kernels, caching, precision optimization, scheduling, and parallelism. As models continue evolving rapidly, which layer of the inference stack do you believe offers the greatest remaining opportunity for dramatic improvements in speed and cost?

For video generation, attention represents one of the largest remaining opportunities because it dominates the compute budget in many models. There’s still significant room to improve lower-precision recipes, sparse attention, and diffusion-specific attention kernels.

Attention is only one part of the opportunity. The largest overall gains will come from combining improvements across the stack. Caching is one example. The KV caching techniques used in LLMs don’t transfer directly, but diffusion models contain spatial and temporal redundancy that can be exploited with the right approach.

Parallelism presents another opportunity. Adding GPUs doesn’t automatically produce a proportional speedup because communication overhead can offset the benefit. We develop tailored kernels that overlap communication with data-independent computation and pipeline it with data-dependent work.

Attention, caching, kernels, parallelism, memory, and scheduling all affect one another. Optimizing only one layer eventually creates a bottleneck somewhere else. The greatest improvements will come from treating the stack as a complete system designed for diffusion’s computational profile.

With newer video models pushing generation times closer to real time, how close are we to genuinely interactive generative video, and what technical barriers still need to be overcome before real-time video generation becomes commonplace?

We’re already reaching interactive speeds for certain applications. MachGen generates a five-second Vidu Q3 Turbo video at 720p in about six seconds and at 540p in under three. That’s fast enough for rapid creative iteration and puts responsive avatars and interactive video within reach.

An example of what I think interactive means. Imagine you’re watching a story, and you can see where the ending is going, and you don’t like it. Instead of sitting there passively, you redirect it: those two characters should go find out what the third one is hiding, and confront him rather than letting it play out. Content you steer instead of content you receive. That’s what fast, cheap generation makes possible, and it changes almost everything we consume.

Getting there commonly requires more than one strong latency benchmark. The whole experience has to stay responsive under production traffic while holding visual quality, frame-to-frame consistency, and predictable cost. Longer videos, higher resolutions, and concurrent requests all raise the infrastructure burden, and the system still has to provision capacity, route requests, and recover from hardware failures without the user noticing.

It’ll arrive use case by use case. Short clips, avatars, gaming, and creative tools are likely to go first because generation measured in seconds is already enough for them. As model quality and inference performance improve together, more applications will cross that threshold.

As AI agents increasingly generate images and videos autonomously rather than waiting for a human to press a button, how does that change infrastructure requirements? Do agent-driven workloads create fundamentally different demands around latency, concurrency, cost, and reliability?

An agent turns a single user request into dozens or hundreds of generation jobs. It creates several options, evaluates them, revises the prompt, and repeats the process, with no human input between steps.

That makes latency, concurrency, cost, and reliability much more important. If each generation takes minutes, the agent’s workflow is too slow to be useful. If each attempt is expensive, autonomous iteration becomes economically impractical. The system also needs to handle many simultaneous requests reliably because one failure can interrupt the entire chain of work.

This is where capabilities such as GPU provisioning, autoscaling, and routing matter as much as model-level optimization. Agent demand is far more bursty than demand from a creative application where a person is clicking a button.

The relevant unit of work is no longer a single image or video. It’s the full loop of generating, evaluating, and refining content. The infrastructure has to make that entire loop fast, affordable, and reliable.

There is intense competition among GPU providers, inference clouds, model developers, and optimization platforms. As hardware itself becomes faster, why will companies still need a specialized inference layer like MachGen rather than relying on improvements from NVIDIA, AMD, cloud providers, or the model developers themselves?

Faster hardware raises the ceiling. Software determines how much of that ceiling will be reached.

We saw this play out with LLMs. New accelerators improved raw performance every generation, and continuous batching, KV-cache management, speculative decoding, and specialized kernels were still what got the industry to production-level throughput and economics. None of that came from hardware.

Continuously optimizing the full production path for image and video generation is a different job from what hardware vendors, cloud providers, and model developers are doing, and it isn’t a core priority for any of them.

There are a few approaches to that job. One is the marketplace model, where models are aggregated and route requests. We don’t think that’s defensible over the long term, because there isn’t control over the layer where the performance lives. We chose to build the stack ourselves and host natively, which is the only way to achieve the latency and cost numbers we see.

That means tuning attention, caching, kernels, precision, memory, and parallelism per model and per accelerator, and supporting managed APIs, dedicated cloud, and self-hosted deployment. As the number of models and hardware options increases, so does the complexity. Our role is to absorb it so our customers can spend their time on what differentiates their products.

You’ve described world models as part of MachGen’s longer-term vision, with many of their computational characteristics resembling diffusion workloads. What does the infrastructure for production-scale world models need to look like, and what applications become possible if generating and simulating visual environments eventually becomes as inexpensive and responsive as generating text is today?

One way to think about our platform is similar to a general-purpose LLM. The same model drafts a legal agreement and answers a question about restaurants. For us, the same stack serves video avatar companies, AI advertising, story and microdrama generation, and eventually world models. Different verticals, one set of underlying performance problems.

World models aren’t our core business today, but they’re what we’re building toward, and we’re actively working on them. Production-scale world models will need very high throughput and very low latency because they continuously generate and update an environment rather than producing a single output. They also need spatial and temporal consistency, the ability to respond to actions, and often the ability to simulate multiple possible outcomes simultaneously. That creates hard problems in memory, data movement, parallelism, and reliability. A world model driving a robot or a game can’t take minutes to produce the next state. Performance decides whether these systems are usable at all.

Computationally, today’s world models look a lot like diffusion models, so the stack we’re building now is the natural foundation. There’s no mature production-grade serving layer for them yet, comparable to what exists for LLMs. We intend to build it.

If simulation gets as responsive and affordable as text generation is now, robots can rehearse rare situations before encountering them, games can generate environments that respond to individual players, and designers can test dozens of possibilities before building in the physical world. Diffusion is where we earn our keep today. World models are our North Star.

Thank you for the great interview, readers who wish to learn more should visit MachGen AI.

Antoine is a visionary leader and founding partner of Unite.AI, driven by an unwavering passion for shaping and promoting the future of AI and robotics. A serial entrepreneur, he believes that AI will be as disruptive to society as electricity, and is often caught raving about the potential of disruptive technologies and AGI.

As a futurist, he is dedicated to exploring how these innovations will shape our world. In addition, he is the founder of Securities.io, a platform focused on investing in cutting-edge technologies that are redefining the future and reshaping entire sectors.