LLM Gateway: One API Key for Every LLM
Listen Now
About This Episode
In this episode of DevNTell, Narb welcomes back Ismail Galloo and Luca Steeb, co-founders of LLM Gateway. They discuss the evolution of LLM Gateway over the past year from an open-source passion project into an enterprise-ready AI infrastructure platform. The founders detail new sub-products like DevPass, AirSide, and The Lounge, achieving SOC 2 Type II compliance, and how AI-assisted and agentic development workflows have allowed a lean two-person team to compete with multi-million dollar corporations.
Key Takeaways
LLM Gateway serves as a unified API layer to route requests across multiple LLM providers, track usage and costs, and manage spend limits in one location.
The team has expanded its product suite with DevPass (subscription-based access for coding agents), AirSide (a marketplace for LLM inference providers), and The Lounge (a multi-modal playground).
By adopting advanced AI-assisted coding and agentic review workflows (such as CodeRabbit, Cursor, and custom agent skills), a two-person team has efficiently built and scaled enterprise-grade infrastructure.
Achieving SOC 2 Type II compliance enabled LLM Gateway to onboard large enterprise clients seeking unified AI spending and custom dynamic fallback routing.
Featured Guests
Ismail
Founder @ LLM Gateway
Luca
Founder @ LLM Gateway
Timestamps(click to jump)
Episode Transcript
Read full transcriptHide transcript
GM GM! Welcome to what's going to be another fantastic episode of DevNTell. So if you didn't know, DevNTell is a 30-minute podcast held every week allowing founders, hackers, and anyone in between the opportunity to come on the show and showcase what they've built. And today I'm really ecstatic to welcome back Ismail Galloo and Luca Steeb, who are the co-founders of LLM Gateway. So if you don't know, LLM Gateway is an open source API gateway for large language models, allowing you to route requests to multiple providers, manage API keys, and optimize costs. So stick around for today's episode, you'll get to meet Ismail and Luca, learn about LLM Gateway, where they've come since a year ago where they appeared on the show, and how you can get started using it today. All right, let's get into it.
Hello, hello. Welcome back to the show, boys. I'm ecstatic to have you back on.
Yeah, thanks for having us.
Hello.
Of course, of course, my pleasure. I'm really keen on folks getting re-introduced or introduced to you guys, because your story is kind of core to why I started this podcast actually in the first place, just kind of seeing folks start with a passion project and turning it into a success like you guys have. But before we get into that, would you guys like to give an introduction about yourselves?
Yeah, let me start. So I'm Luca. My whole life I was a software developer working on different projects for different companies. I was also always interested in building my own stuff. And I've known Ismail for a long time, I think almost a decade now, and we've collaborated on things a few times in real life. Around one and a half years ago, we started LLM Gateway because we were working on AI stuff, on automations, and personally I'm more into DevOps. We realized that AI models even back then were very split into different platforms and very fragmented. Every API has different integrations, you need to sign up for different things, different billing, and then some of the platforms don't report any usage insights and you have to wait for a day. That's why we kind of decided there must be a better way to do this, and that started LLM Gateway.
Ismail, you want to introduce yourself?
Yeah, sure. I'm Ismail, also known as Smakosh on X and everywhere else. As Luca said, we know each other for almost a decade now. I started as a designer, then turned into developer, but nowadays I label myself as just a builder. I'm also working full-time in my job at The Token Factory. I call it The Token Factory because most developers know us as Prompt10 and consuming tokens. Yeah, I had the same idea to provide one API to connect all these providers and models. And when Luca reached out on Telegram, asking if I was interested to join and build this, my answer was yes, and we immediately started building. There was no planning, no meetings, no product strategy or market strategy. We just went on and started building. We were using back in the time Cursor and v0 where we started vibe coding, but we were still manually checking the codebase, doing smaller PRs because back then AI was not really good at producing code. But over time, when Opus 4.5 came out in December, we started trusting the code it produced. But even though, we were still verifying and reviewing some code. But lately, especially when Fable 5 and later on Astra came out, we don't read the code anymore, but we're somehow doing this agentic coding. We have another AI that reviews the code every PR. We're using CodeRabbit; it's free if you have an open source product. But in our workflow, we generate the code, we test it, like we have AI test it, it records the recording, we have a bunch of skills. We pretty much made code really good for AI to produce less slop in the future.
Yeah, amazing fellows. And yeah, indeed a year, it's quite interesting to see or reflect on how much AI has grown in terms of capability, right? I remember back in the day—a year ago—I was wrong with Cursor too, I was hesitant on like, 'Oh yeah, okay, this is cool, I can get rid of some boilerplate stuff for me.' And then like you said, when Opus 4.5 came out in December, I was like, 'Oh, this is now pretty good!' So yeah, I think it's just going to keep going unless external forces act against it. But that's another podcast. Coming back to LLM Gateway, for folks who might not have caught the previous episode, do you guys want to give an overview of what LLM Gateway is?
I think the easiest explanation is that it's one API to use any model with any provider. But you can still track all the activities, like what happened when you prompt, the cost, the usage. But it also comes with some sub-products. We now have DevPass, which is a subscription-based product where you can use it with any coding agent. And we have AirSide as well, which is an LLM provider marketplace. Because of the rise of open-source models, there is a lot of inference companies doing inference for these models, so we grouped them in one place, and they compete who gets more traffic by lowering the price and doing inference really good. We also have The Lounge, which is a chat playground and an image, video, audio studio in one place. You can still use your credits there, or you can subscribe to one of the plans to use it. Yeah.
Awesome. And yeah, DevPass, Lounge, these are just some of the things you guys have introduced. And I know you guys have crossed some pretty big milestones as well, so I'll give you a chance to talk around that as well, and kind of the evolution you guys have gone through from a year to now. We'd be interested to hear.
I think one of the things that Luca said last time regarding the roadmap was adding video, image, and audio models. I think Luca can elaborate on that.
Yeah, I mean, I think we have expanded the product in every detail imaginable, I would say. I think for all of these kind of things we have added those features. I think the shift, maybe worth mentioning—let me know if you want to step back first—but worth mentioning is that we have bigger companies interested in the Gateway. So I think our last podcast was quite a long time ago, and we were still super early back then. So we have lots more people that signed up and buy some credits and provide feedback. And there's always a bunch of features that you need to catch up with competitors, and also fixing bugs takes time. Adding new features which are specific to a provider, which only that provider supports, but we still want to support at the gateway level, right? But what's interesting, I think—and we can even go deeper into that—interesting is bigger companies are starting to use LLM Gateway. We also got the SOC 2 Type II compliance, which bigger US companies need. So that allows us to actually sell to bigger customers, specifically in the US. And what they need is... they have the same problem, right? They want to use Claude Code, Codex, they spend a lot of money on AI, they have contracts with these providers directly, but it's still fragmented. Even if you have 2 or 3, just for the contracts perspective, you have to commit to specific spending amounts per year. And your developers, some may use Claude Code, other people might use Codex, and then how do you set limits, how do you ensure people don't spend too much money on tokens in the end? And at the same time, companies also want to use open-source models; however, they can't easily do that because they have contracts with flagship providers, right? So why not have one API where you can use everything, right? So we have a lot more interest recently from big companies that want to unify their AI spend. And it's quite interesting because some of these companies spend tens, hundreds of thousands every month, maybe even millions. And they want to reduce cost, right? So the open-source models are catching up... and yeah, so they can use LLM Gateway, then it doesn't matter if you use Claude Code or Codex, you just put two config lines and you can use any model, right? And then managers can set spend limits in one single place, and it's almost necessary at this point, I would say.
Yeah, other than the models and providers, new APIs we supported, we also improved our routing, real-time routing, cost controls. We now have automatic retries and fallback. We added support to regional models, for example Azure, GCP, Alibaba, they have different regions so you could use the same model in a different region. We also added a lot of cost and speed strategies, dynamic routes, which was a neat feature requested by one of our enterprise clients. They wanted a way to build their own custom dynamic router so if this model fails, I would like to fall back to this other model. And they can do it through the API or through the dashboard with a GUI. We improved a lot the Gateway in terms of latency, caching, and all of that. There is a third-party company called ComputeSDK; they do this weekly benchmark. We actually ranked first on August 7th. And if you go check who we are competing with, it's multi-million, billion-dollar competitors. We're only two guys behind this LLM Gateway. The other things that we added since we last spoke: team access, SSO, audit logs, guardrails, provider compliance policies so you can define if you want to use this model but from providers that comply with GDPR or ISO 27001. You can even specify the headquarters of the company providing the model. As you can see, these features belong to the enterprise edition of the Gateway. We've been working really close to our enterprise clients. Whenever they request a feature, there is no meeting, no planning between me and Luca. We immediately go, one of us takes it, owns it, ships it, and communicates it to the enterprise client. And the quality is really high, it's as if 20 or even more members worked on it, and this is because of how good the recent models have become.
I think one thing to add there is where we can really shine as a small team. When a big company works with another big company, if they request something, it's usually a long process. That's why some teams prefer the smaller team, which was not very realistic 1 or 2 years ago because 'can these guys actually have the ability to ship this fast?' Now it is, right? So now we can stay so flexible and ship things really quickly and in high quality as a small team.
Yeah, for the developer community, because we're an open source LLM Gateway, we did ship a lot for developers since we last spoke. For example, we have some agent templates that they can start with. They are already using the AI SDK provider from us. We maintain our AI SDK provider, recently upgraded to the latest one. We have now a CLI that your agent can use, you don't have to go to the dashboard to do things. We upgraded our MCP; you can get your usage and all of that through MCP as well. We forked OpenCode and got rid of all the providers, kept only ours, and fixed some billing issues because OpenCode does not track sub-agents, how much they spend. So we fixed that so billing is accurate. But we still work with OpenCode by providing our provider within their website (models.dev) where we keep in sync. So when we add a new model, if you're using OpenCode, you get the latest ones showing up there. Overall, we're still open source. The Gateway can still be used for hobby projects and non-commercial use. Anyone can check the code and get to know how we're building things; they're free to explore that.
Yeah, amazing. You guys have jam-packed the product with a ton of useful features, for sure. And from the provider perspective, I know you guys probably got slammed at some point, and that led you to building the AirSide product. Did you want to speak on that?
Oh yeah, AirSide. Smaller providers used to reach out through email to me or Luca, we'd get them added to our Discord server or Slack workspace. They would provide that they support these models, we had to manually create a pull request, test their models, verify, merge, and then they are live. But for every change or new model they add, we have to do these PRs again. So we thought maybe it makes sense to self-serve them through a smaller product. That's when we came up with AirSide. We tried to keep the branding towards airports, passports, and traveling. So AirSide is like an airport, and an airport has airplanes or airlines—those are the providers. And they compete between each other to drive the airport tax, which is the gateway margin. The more they increase the gateway margin, the more traffic they get, but that depends if their model throughput and availability is better. They also can lower the price to the end user, and that's how they get more traffic. They mostly get their traffic from our other sub-product, which is DevPass. The difference between LLM Gateway and DevPass is in LLM Gateway you buy credits, they don't expire, you can choose the provider you want to use for a specific model. But for DevPass, you subscribe, you get extra usage, but you cannot pin a specific provider. But you can still play with the routing strategy—whether you want the cheapest provider, automatic routing, or more available provider. After a lot of our subscribers were asking for zero data retention feature, we allowed them to toggle 'no AI training', so we don't route to providers that train from their prompts. So that's where the traffic comes to these AirSide providers. We started sharing our usage on a public page, I can send it in the chat. We have processed almost 2-3 trillion tokens so far through the Gateway, and we report daily on another page—daily usage, weekly, and monthly. If you go to that link, you'll notice that the open-source models are the most used ones.
Yeah, there's been quite a renaissance on the open source side of things, and you guys caught on the trend early as you started open source, and that movement has exploded this year especially. Why did you guys start open source, why is it so important for it to be part of the DNA?
That's a fun story. So we were looking at competitors back then, specifically OpenRouter, which is so well known. It's called OpenRouter, but there's nothing open about it, because it's just a closed-source platform. So we were thinking that's kind of funny. And then we decided maybe it's time for a fully open-source gateway. There are other solutions like proxies, but they have more specific use cases, and what you want to have is something like OpenRouter, but still open source. The long-term goal is that enterprises can use it, install it into their own infrastructure, fully self-hosted for internal use or white-label in the long run. If people want to self-host it, then open source is mandatory, right? But what we still saw a lot recently is that enterprises still end up using our cloud version because of maintenance. But it still builds that trust that people can verify things, and we have a community that contributes, helps us fix issues and pricing changes, so it's a whole different dynamic. It's more welcoming, it builds trust, and ultimately it helps companies to deploy it to their own infrastructure if they want to do that.
Absolutely, absolutely. I think this trend is going to continue, especially as more and more closed-source providers on the model side continue to get more strain on the compute side, things get more expensive, token usage goes down. It's going to be interesting to see how it unfolds. In terms of developers wanting to get started, what's the best way for folks to get up and running?
Since we work with enterprise, we have different products depending on the user. For example, if you work in a big company, your company most likely is using an AI gateway, whether another competitor or Copilot. LLM Gateway makes sense because you can use LLM Gateway as an API or with coding agents. With Copilot you can only use it as a coding agent. The good news is you can still use Copilot coding agent with LLM Gateway by bringing your own keys. But if you're a smaller startup or an individual/solo developer, DevPass makes more sense for coding if you want to use all the coding models with all the coding agents. If you're a consumer, then The Lounge makes more sense; it's a chat playground, image/audio studio in one place.
Definitely. And I was looking at your docs and the usage for DevPass is quite different. It seems like for every $1 you get $3 of usage. Did you want to speak on that?
So because of AirSide, we have a minimum percentage for the gateway margin. That's how we can afford to give extra usage. But there is a change coming in October 15th when we'll lower the 3x to 2x. It was mostly due to abusers using the product wrongly. But we did tighten up things and we have now a set of coding agents that we allow to use with DevPass. We don't allow resellers; there has been a lot of fraud, relayers, and resellers. We've done a lot to combat that.
Gotcha. Yeah, there's like a whole black market of people stealing credit cards and using it to buy tokens and not getting blocked.
That's actually worth mentioning because everyone, including us, underestimates how bad that situation is. Fraud has always been around, but if you sell SaaS, it's not a big deal because it's virtual and doesn't directly incur any cost. But for tokens it is, because you pay us $100, use the tokens, and we pay the provider directly $100. If people abuse our service, we have incurred the cost already, and if it's stolen credit cards, we get a chargeback from Stripe. So then we have a loss. All AI gateways are dealing with that. We have a little bit more risk because we are not VC funded. Some people can just hire a bunch of people who do fraud detection all day. It's a cat-and-mouse game; you cannot prevent this from happening, but you can have checks, rate limits. It's a big deal and we're working on this.
Yeah, absolutely. Once you start growing, the bad actors start popping up. Moving on to the founder side of things: you mentioned it's just you two and a bunch of AI agents managing this company. For founders who are watching today who might be on the fence of starting their own company, what advice would you share with folks watching today?
Personally for me, I've always been doing projects on my own at the side, and the biggest problem was you can only do so much. Two or three years ago, pre-AI, you have the knowledge to do things, but it takes so much time to code all of that up, even just a dashboard like AirSide. Anyone could build it, but it takes a long time because there's lots of features, wiring, making sure it's robust, tests, authentication. Previously it would be just too much, that's why people have funding and hire 10-30 people. With AI it's changing a lot, and in the last year the models are getting so good that they can fully handle things end-to-end, fix it, show you a demo video of what it wired up. Everyone can basically build a real company now. That creates a lot of competition, so you still have to think about marketing and distribution, which is now way more important. But you don't have to raise millions of dollars to hire a team—you wire it up, prototype it, and if it's not good, try again. It's shining for individuals and small teams.
My best advice would be: just build it. Because if you got this idea and you're still thinking about it, somebody else out there has already shipped it the next day. Whoever gets to distribution faster is the winner. I think it's funny, it all started because of 'Attention Is All You Need', which is the popular paper, but it also applies to this vision: whoever ships first gets the attention and converts the attention to paying users. The only thing that a lot of people make mistakes about is scaling and infrastructure. If you spend time on engineering to make things scalable and handle a load of traffic, then you will not get to distribution. But if you get to distribution and your product is crap, nobody would want to pay. So you need to play in between these angles and iterate. For us, Luca is the infrastructure guy—he makes sure the infra can take as much as possible (we've handled 82 million requests so far). I'm the distribution guy, and we try to balance things. For AirSide, it was built in one day, launched, got one provider to test it and provide feedback, implemented fixes, got more onboarded, and then Luca stepped in and added end-to-end tests so providers would run them before being listed. To summarize: don't spend a lot of time on engineering, don't spend a lot of time on distribution—try to balance it in between, iterate over staying on one thing. I would also recommend that you target enterprise rather than play with a consumer audience.
Excellent advice from both of you. I really appreciate you guys going into that. Hopefully somebody watching/listening today, this will be enough to push you to start that passion project you've had in mind. I know we're at time, so I want to give you guys a chance to share how folks can follow along with you, the latest updates, and how to reach out.
For latest updates, they can check the changelog. We categorize it based on products. They can follow each product update or overall updates. We're also up to date on X. They can join our Discord server if they have any feedback or things they wish we have.
Yeah, on Discord people join and provide feedback, and we fix issues very fast. If anyone has suggestions that are a great idea, we just add it because it's so easy to add if it makes sense.
Excellent, excellent! All that information will be in the description of the podcast below. So if you're interested, definitely give LLM Gateway a test and get involved. With that, Ismail, Luca, thank you so much for taking the time out of your busy day to come chat with us today, and wishing you guys all the best as you continue to grow and expand here.
Yeah, thanks again for having us.
My pleasure, anytime. And with that, just want to wish everybody a very happy Friday, happy weekend wherever you may be, and we will catch you back here next week for another great episode of DevNTell. Until then, have a good one, folks. Cheers!
Cheers!
Listen On
Resources & Links
Share This Episode
Share on XWatch Episodes Live!
Subscribe to our event calendar and never miss a live episode.
View Event Calendar