Intelligence for Anyone, Anywhere: A Look Into Gradient’s Mission to Decentralize AGI
Listen Now
About This Episode
In this episode of DevNTell, host Narb interviews Alex Mirran, the Business Development lead at Gradient. They discuss Gradient's mission as an AI research and development lab dedicated to creating open intelligence through fully decentralized infrastructure. Alex shares his background in distributed systems, crypto, and AI, explaining how his experience in Filecoin and Ethereum communities led him to Gradient. They explore current trends in the intersection of AI and blockchain, emphasizing the importance of specialized AI solutions for enterprise needs. Alex demos Gradient's product, 'Commonstack,' which allows users to access various frontier AI models in one place, with features like smart routing to optimize costs and ensure reliability. They also discuss 'UncommonRoute,' an open-source local router for LLMs, and touch upon Gradient's origin and future focus on creating added value for AI customers.
Key Takeaways
Gradient aims to decentralize Artificial General Intelligence (AGI) to provide open intelligence through a fully decentralized infrastructure.
Smart routing in AI systems can lead to substantial cost savings (up to 82%) by selecting the most efficient model for specific tasks.
The intersection of AI and blockchain is still in its early stages, with potential in areas like compute aggregation, verification of outputs, and agent payments.
Enterprise customers require highly specific AI solutions, moving away from one-size-fits-all models toward more verticalized and niche integrations.
Open-source AI models are rapidly catching up to closed-source frontier models in performance, especially through optimizations for consumer-grade hardware.
Featured Guest
Alex Mirran
BD Lead @ Gradient
Timestamps(click to jump)
Episode Transcript
Read full transcriptHide transcript
GM GM, welcome to what's going to be another fantastic episode of DevNTell. So if you didn't know, DevNTell is a 30-minute podcast held every week allowing founders, hackers, and anyone in between an opportunity to come on the show and showcase what they built. And today I'm ecstatic to welcome my good friend Alex Mirran, who is the Business Development lead at Gradient. So if you don't know, Gradient is a AI research and development lab dedicated to building open intelligence through a fully decentralized infrastructure. So if you stick around for today's episode, you'll learn all about Gradient, get to meet Alex, and learn how you can get started using Gradient and some of their products today. Alright, let's do it.
GM GM, Alex, what up buddy? Pleasure to have you on the show man, long time coming. Alex Mirran: Hey Narb. Man, I have to say it's been such a pleasure watching DevNTell grow. I'm super honored to be here, super excited for our chat. Narb: Thanks buddy. Yeah, appreciate it man. It's been an interesting journey on the podcast itself and yeah, I'm ecstatic to have you on today. I know you and the folks at Gradient are doing some amazing work in the decentralized AI space and AI space in general, so really keen on learning and having our audience learn as well. But before we get into that, would you just like to give an intro for folks who might not be familiar with yourself?
Yeah sure. Hey everybody, I'm Alex, lead the BD team over mostly in North America for Gradient. And yeah, I've been working in distributed systems for quite a long time. Worked with Narb previously at a previous company, also in sort of the distributed AI world. Before that, built a lot of stuff in like the Filecoin IPFS community, where I was mining Ethereum when it was proof of work. So have really been in these sort of intersections of distributed systems and crypto and over the last maybe four or five years, most recently in AI. And so yeah, focused on a lot of things like inference-related. But yeah, connect with me, I love to learn from people who are much smarter than me. And so Narb for example, that's why we became friends. But anyway, yeah, super excited to talk more.
Excellent, excellent. Yeah man, and yeah, I mean it's an understatement you being into AI. You're quite involved in the AI space now. I see you hitting up meetups in Texas and hosting some of your own. I think you guys or you yourself were at something called Closton or something yesterday or this week. Did you just want to briefly give a talk around that?
Yeah, yeah, big shout out to the Closton team. They organized an event with ClawCon, sort of the global open claw group last month. So this month they had their next event. And yeah, great turnout. The whole AI scene, I feel like in sort of like the Bay and New York and Austin are kind of developing very quickly. And like when I go to these events, it's incredible. We hang out with people who are like a brand new startup that's just barely got an MVP or even just an idea. All the way to yesterday, talking to people from like Dell and people from Meta show up and people from Google show up. Actually there were some people from Google at the event. So yeah, it's just really cool to see broad buy-in on a technology from everyone of every size and every shape. I think everyone's trying to figure out their strategy and sort of wanting to see what people are up to. So yeah, Closton was a cool event yesterday. I was able to demo some of the stuff we're working on and there were some awesome other groups there demoing. So yeah, big fan of those guys. And there's a bunch of awesome groups here in town, like AI TX does a lot of good stuff, Hack AI does too. So yeah, I'm happy to just kind of jump around to those groups and support them where I can.
Excellent. Yeah, I love it man. Yeah, getting involved is like the best way. Building from home is a good start. Getting involved in the community and seeing what other people are doing or building things together is really taking it to the next level. Alex Mirran: Yeah, I mean I definitely spent at least a few days preparing for my like five-minute demo, right? Like you said, you got to build, you got to figure it out. You don't have anything useful to say unless you're actually trying stuff out. And so, that yeah, that's an understatement about building and then going and sharing. Narb: Of course, oh yeah. Yeah, people don't realize how long it takes to actually make that five or ten-minute video. Spend maybe a couple hours, few hours, it's a process. But you mentioned it earlier, a lot of very, very interesting developments are happening in the AI space. I mean, we're only April, we got Claude Opus 4.7 yesterday, we got Qwen 3.6, we got Gemma 4 in the last couple weeks. So a lot of strides are being made on the open-source and closed-source, I guess, model frontier. We don't see quite a bit of chatter on the blockchain and AI front as much anymore. Maybe it's getting hit in the dust or people are building in silence. I guess you being at the forefront of this, I guess what kind of driving themes have you seen thus far this year around that intersection?
Yeah, I mean I agree with you, I think it's just really early days and I think everyone's trying to understand product-market fit for certain use cases. So I think probably the most original one that still has huge amounts of value is sort of like aggregating compute and being able to sort of incentivize that compute to be available, cool things around like verification of outputs, you know, running in zero-trust environments, so running in encrypted containers on hardware that are like encrypted enclaves to where the person running the machine doesn't know what the user is actually doing for privacy. So yeah, there's a lot of cool stuff there that I think there's a lot of teams working on. I think EigenLayer, I think you and I were talking Narb, they had a sort of a new I think it's called Darkroom or something like that, but it's like aggregating Mac minis and they're putting idle compute to work. So yeah, that's pretty cool. I do think some of the stuff I think we learned together working together on this was these types of compute networks tend to be one-size-fits-all. And so I think some teams that are working on structures that are not necessarily locking in customers to one specific architecture and one specific security choice, one specific geography or one specific, like every customer, especially on the enterprise side, has very specific requirements that like a standard out-of-the-box DePIN network needs to adapt to. So I think all the DePIN networks are figuring out how to do that in a way. Mostly it comes from building layers on top, like Filecoin's building layers on top and then like Bittensor ecosystem, they have subnets. I'll talk more about it, but at Gradient we're thinking similarly in that world where clients could potentially spin up their own sort of subnet or cluster to do training and then spin it down. So I think that's the first thing, compute. The second thing, agent payments seems to be a big discussion. I think that is extremely early days. Like we have MPPP from Tempo, we have XRO2 from Coinbase I believe, or maybe it was a collaboration. Those protocols for basically kind of standardizing how agents interact with commerce using crypto rails like stablecoins. It seems very interesting and the concept of using smart contracts to sort of program what your agent can and can't do is very interesting in terms of money. I don't know, like what do you think Narb? I mean, I know you build a lot of automated systems. Like how confident are you in giving your AI agent a wallet to sort of achieve a goal that you've set to it? I don't know, I'm getting there, but I still think people are trying to figure it out.
Yeah, personally I still don't think it's quite there. I still don't have that trust level set for an agent to do right operations against the wallet. Read, yeah sure, you can go nuts. But handling money, I still am on the fence. I think it'll get there, maybe this year, maybe next year, who knows. But it's kind of I've seen the same themes kind of being built out here around like the payments. I think eventually I think I've told a couple people this already but we always complained, us in this crypto space about how difficult it's been to use crypto wallets in general, the UX around it. Ironically everything kind of that's in place now for being able to create a transaction, sign a transaction, these things are very basically created for AI. It almost seems like all the information is there for an AI to be able to discern what a transaction is, how to perform it. So I think all the building blocks are there and yeah, I'm really excited to see what builders out there kind of do to refine that and kind of make that a reality. I mean agents eventually will need to buy things, maybe they'll buy things from other agents, and I think crypto's the perfect rails to do that. Ninja magic internet money, right?
Well that's what I was going to say is yeah, AI systems working with internet-native currency, like that just makes sense. Global payment rails, instant settlement, like these are bear assets that have final settlement characteristics. It makes sense. I think the use case that seems to be the most popular in this world right now would be like trading, so creating trading bots that are good. You know, we've been experimenting with doing some post-training on open-source models to sort of outperform like using a frontier model to trade. But yeah, it's early days for that kind of stuff. But it is exciting. And then the last thing that I feel like is really cool too is sort of provenance trails for data that gets used for training. So there's some cool projects, some cool companies that kind of maybe started out in the data labeling world and now they're using crypto to incentivize the collection of more data. It's kind of a standard business model, you know, you pay someone to either label something or to provide subject matter expertise on something and then the company will structure that data in such a way that they can box it up and sell it to frontier labs and such. So I mean that's a really straightforward business model and I think using crypto to coordinate that, incentivize that, you know, Narb you and I have worked with some decentralized protocols with regards to storage and like verification of data like IPFS and such. Seems like a really obvious place to use that kind of tech but you know, I think it's just yeah, it's day zero.
Yeah, I mean at first glance you might be like ah like we're all this way in 2026 and it's day zero but at the same time the other side of the coin is oh there's so much ocean of opportunity there for people to make their own businesses around it or whatnot. I think it's very exciting, very exciting time to be alive so to speak.
Hell yeah, I mean it's as simple as there's people that I know here in town in Austin that they just they have ten or twelve clients and they're just implementing open claw systems for them or some kind of business automation. Like N8N was used to be kind of the thing that people were asking for, now it's more like agentic harnesses like Hermes or OpenClaw or OpenCode, stuff like that. And yeah they're making like five, ten, twenty, fifty thousand dollars per client and it's very much feeling like sort of the when everyone needed a website in the nineties or everyone you know now needs to understand how to like implement these things into their systems. And I will say too I feel like the groups like the consulting doing integrations is really, really cool but what's even cooler is productizing for specific verticals. So like we talk to a lot of companies that use inference because that's one of our main things that we sell and the companies that are growing like crazy are ones that are very niche and verticalized and they know the problems of everyone in their sector and they just know all their integrations like their ERPs, their CRMs, you know all their internal processes, they know their problems. And they're just like hey we have an agent deployment platform, we can solve X, Y, and Z problem, we did it for a hundred other service businesses. That's where I'm seeing like huge amounts of growth very quickly.
Yeah, it's amazing to see the velocity some of these teams are working at and one of these teams I believe is Gradient, I don't think you'll have any argument around that. Which is a nice segue to yeah, what is Gradient exactly? What do you guys do?
Yeah, yeah, yeah, so we sort of originated as a research lab. Our team has very deep expertise in distributed systems and we have some really strong AI researchers. And so the lab was focused on tools to open up access and provide the low-cost capability to people to run inference and reinforcement learning. And so we have a bunch of open-source tools, Parallax being one of them, is similar to sort of like an EXO, you can run it, it's open-source, it's free, you can run it on your computer, cluster together Mac minis or Apple silicon and NVIDIA hardware. That became part of Echo, which is our distributed reinforcement learning tool. And then Commonstack is our sort of OpenRouter style cloud. So we aggregate inference APIs from hosting providers like Novita and others and we have all the frontier models open-source and close-source. And we have really good rates, really good volume discounts on that, so hit me up and happy to talk more about that. But yeah, our mission generally is like use distributed systems to open up access to these very taxing, complex, expensive systems so that more and more people can join in on the AI revolution. Our tools, I'm really impressed, the team is incredibly good at abstracting away the complexity of even just running inference. I mean like Narb, you spun up inference APIs like I've seen you do it very quickly but it's incredibly complex to orchestrate parallel processing, right? And so to abstract that away and then you have like a simple UI and it's all free running locally, like that's really cool.
Yeah, people don't realize because there's so many options out these days people think it's very easy to spin up and do. It's actually not, it's very the coordination like you mentioned is yeah, no cakewalk. Alex Mirran: Yeah, I mean managing an API by itself is a job, right? Is a task, right? To vertically, horizontally scale it, blah blah blah. But then to manage the underlying infrastructure, the hosting of the AI models, that's a whole other game, right? It's yeah, I mean I don't know what yeah, I guess it's interesting how I think some of the biggest labs are creating their own custom inference engines. And so like the XAIs, the Anthropics, the OpenAIs, the guys with the billions are spending the time to build their own software stacks. So that's pretty interesting, but that's more of a proprietary thing I think. Narb: Yeah, agreed, agreed. Yeah, basically like whoever owns the compute will own the world so to speak. I mean like every company is kind of running running towards that end goal it seems. But I guess before we kind of deep dive a little bit more into the different product offerings of Gradient, be interested to hear if you know anything about kind of the origin story behind Gradient and how and why it kind of came about.
Yeah for sure. So our CEO Eric came from Sequoia Capital China. He also co-founded a company, a live streaming company and exited from that in sort of in the Web3 world as well. And then our other co-founder Yuan came from Helium Network. So he was you know sort of on the original team leading the growth for Helium. And so yeah they brought together people with sort of that world of exposure and expertise and we actually started out before I joined building a decentralized CDN. So we had an enormous CDN, I think it was like 900,000 nodes or something like that globally. And so yeah, so that was sort of I think the initial place where Gradient started. And then over time we just saw more and more opportunities in you know applying the distributed systems expertise at the company to specific AI processes like inference and training, reinforcement learning specifically. So yeah, so the company sort of redirected energy and most of the engineering capacity to the we've put out like four or five papers and so we started looking at which one of those research endeavors are we going to commercialize. And then the engineering team jumped onto that. So yeah, that's kind of how we landed up where we're at right now. But we do have so it's interesting because we do have very deep blockchain ethos and the concept of sort of owning your own things and open-source, local-first, like all of that stuff flows in our products. And then but like right now we are very focused on people with AI problems. And so over time like we're looking very carefully at where we can implement blockchain functionality into the stack to add value to our current customers and then to help us get more.
Yeah, I mean makes total sense to me. You want to make money first and then you kind of expand out from that. But yeah, as I understand Gradient is made up of a bunch of different technologies, open-source like Parallax, Echo, Lattice. I don't think we'll have time to get into all of these, but since we are a primarily developer-heavy audience, you mentioned before Commonstack, your primary AI inference layer for people to be accessed to compute there. I guess did you want to kind of take us on a tour of that and possibly some other things?
Yeah, yeah, for sure. Let me pull it up. All right, can you see that? Narb: You're good to go. Alex Mirran: Okay awesome. So yeah, it's a really similar product to sort of when you look at different routers, right? You're solving a couple different problems. You're allowing customers to have all the models they need in one place. So that's the first thing, is we have all of the frontier models coming out of, I mean Opus 4.6 you know came online yesterday, Qwen 3.6, Gemma 4, all these models are here. And we actually just proxy to other hosting providers. So we work with like Novita and others in that kind of category of company and have SLAs with them for performance guarantees and security guarantees and that kind of stuff. So that's the first kind of interesting piece. The other thing that's cool about it is you have one API key, one API format, and you just change out the model. So you know, similar to OpenRouter, you can just really easily switch out models in your system, which right now is a huge, huge important piece of the agent puzzle. Because the big discussion is which models are good at performing which tasks and like where can I optimize my costs? Because you know, for example, I mean we can just see we have retail prices, Opus 4.6 is five dollars per million input tokens, 25 out. MiniMax M2.7, which you know admittedly isn't you know nearly at the level of performance but for certain tasks it doesn't matter, you know they charge 30 cents. And it's dollar twenty. Even Qwen 3.6 which is super, super good and super strong, you know dollar forty. And so you know the value of having all the models in one place is that you can build your system to to route to these different ones based on the tasks. And then we also do some cool stuff with our routing, so when we notice a provider giving us latency or giving us 429 rate limiting errors or even it just goes down, we switch you to a different provider of the same model. So we're also kind of helping you preserve your business and make sure that your inference is always on. So yeah, so come check it out, you know you can set up a team and you know delegate keys to people. We obviously have a playground, the docs are really, really developed, you know we have every you know streaming, reasoning, tool calling of all kinds, you know all all the good stuff. So yeah, let you know let me know, hit me up on X or send me an email at alex@gradient.network and happy to like talk more about the details around how this works and then also yeah, we're getting you know upwards of 20% discounts from our hosting providers and so you know for volume token usage we can we can we're passing that on to the the end customers.
So that was the one thing I wanted to show, but I also wanted to show you something fun we built that's open-source on top of sort of you can think of it as kind of on top of Commonstack is UncommonRoute. So along the lines of what I was just saying about people needing to optimize their cost, we built this local router. And we're seeing like 82% cost savings by simply looking at the prompt, identifying let me just explain it here. So we look at these three signals: the conversational structure, the semantic similarity to known task patterns, and then you know feature complexity. And we then assign that to a model and fire the request off. So it's a local proxy that you would run on your run on your computer and you can proxy to any API. I mean presumably you'd need a router of some kind, so like a Commonstack or an OpenRouter or something like that because you're going to be switching models, you don't you know you need more than one model. But yeah this is cool, this is what I demoed yesterday and you know we're really we think this is really high impact. We actually created our own bench too, and so we did a lot of the tests with our benchmarking system Commonstack Bench. But yeah anyway, you know so you can set it for auto, which is like a balanced approach, fast, so cost-first, or best, quality-first. So yeah check that out, I think can you see the URL? I guess maybe you can't. Narb: If you can zoom in a notch or two. Alex Mirran: Oh sure yeah. Narb: Yeah that's good, perfect, perfect. Alex Mirran: Okay awesome. So yeah, so check that out and it runs pretty well. I mean I will say that depends on your network, right? You want to make sure that you have good internet connection because latency can be can be a consideration when you're doing an extra hop. But but yeah, no we're super proud of of how far you know Commonstack has gone and now what we're thinking is you know what do we sort of add without going too crazy, what do we add to provide value on top of the system.
So that when, you know when we talk to customers they're like alright we understand what you're building, you know now we're thinking about alright how do we add value? Because at the end of the day like the models are becoming more and more of a commodity. And also, you know Sequoia Capital put out a really great paper called deliver work results not necessarily tools. So the concept was if you build a tool that uses a model once like Anthropic pushes a new feature or a new model drops your tool could be obsolete. But if instead you're selling tools in functionality and outcomes, so a really simple example is like you pay 10k for TurboTax or whatever, but you would pay 100k to someone to use TurboTax to close the books. And so that 100k is where the opportunity is for automating. And it's so anyway, that's how we think about the product roadmap for Commonstack is like where can we add value on top of the inference? So UncommonRoute is one and then the other exciting thing we have in the pipeline is you can check it out on our website gradient.network. We wrote a paper called Echo and Echo 2 is a distributed RL framework. And so what we're thinking is that we will use the the results are coming back really strong, we're training models on distributed clusters, we're seeing huge cost and time savings. And so we're thinking about integrating that into Commonstack sort of as a continuous learning functionality for your AIs that are using our API. Anyway, more on that stay tuned. Also C-Dents 2 is going to drop probably in the next week or so, so if you want video come check that out. But anyway, you know long story to say a simple thing which is you know we provide the best inference and you know we're looking to build value on top.
Awesome. Yeah that's yeah you guys have certainly built out quite quite an array and it's great to hear you guys are keeping that keeping that steam steamroll going there. And yeah it seems like you guys are really concentrating on the user developer experience there and making a product that people would want to use in and out day in and day out. So just to just to kind of come back there onto the UncommonRoute just so I get my understanding correct, you guys are basically you have a another AI layer before you actually hit the main inference layer to determine what model will perform the task the best correct and and cost optimize for that as well?
So that is UncommonRoute. Yeah sorry sorry if that was naming I even I like get them a little bit twisted. So UncommonRoute is the local proxy you can run and that's the fancy, you know smart routing. Commonstack out of the box doesn't have that right now. We're talking to people and seeing if they you know I would be curious would you use that like if it was in Commonstack would that be because some developers are like I'm choosing the model, I don't want you to choose for me. Fair enough. So yeah Commonstack out of the box doesn't have that but what it does do is it looks at latency and downtime and 429s. So we do check for that on Commonstack and you know that helps you make sure your business you know never loses access to AI.
Gotcha, gotcha. Awesome. Yeah, I mean as a consumer of a product, I always like choices. So for like a power user, yeah, yeah for sure, if you know what you're doing, you choose your model. But what if it's somebody who doesn't really know like what the difference is between like an Opus 4.7 and I don't know, MiniMax 2.7? And they're just trying to get it to write like a paragraph or something like that, right? Then obviously you'd want to use the cheaper model to do that because it's comparatively just as performant. So yeah, I think there's personally there's room for both. But yeah.
Yeah, you know something that's been such a fun conversation with people is like what model do you use for your business and why? Every week the answer changes and I actually have been just keeping a spreadsheet with all the answers that I get. And it's been actually a really cool discussion to have with people who just they're like I don't have time to look at evals, I don't have time to keep track. You know can you just tell me like I need a really good reasoning model that's not too expensive? Or I need a really good like tool calling model or I need a really good coding model, whatever. So that's just such a dynamic conversation. It changes every week, you know, I mean you there's a joke it's like you need to be like unemployed basically to know everything that's going on in AI. Now crypto people, we're terminally online and so you know we're hustlers, we're good with distributed systems, so there's a reason there's a lot of crypto people in AI. But anyway, yeah, the whole eval situation is so interesting and it's yeah, constantly changing.
Yeah, yeah, and I think it'll continue to change. There's just too much happening now. Like AI in itself has allowed people to to ship way more frequently. I think for example the people people at Anthropic I think they're shipping maybe two or three, four, maybe even more than ten times a week or a couple weeks. So it's really hard to to kind of keep up but like these products are evolving like you said way faster than people can keep up. So pro tip, you can use AI to gather all this information for you so you don't have to be terminally online.
Yeah, that's true, that's true. So wait, I want to know your perspective Narb. Like as an engineer, I hear very mixed reviews on how people feel about like using you know Cloud Code for example. Like I've used it for front-ends, I've used it for like a ton of I mean even like setting up APIs or setting up you know interesting things. But doing an entire deployment horizontally scaling it and like making sure it's secure and hardened like I don't know what's your what's your feeling right now about where we are?
I still don't personally I wouldn't rely on it 100% to to do every single thing that I'm like out of my day-to-day. However, it does save me a certainly around like 80% of my time just doing like regular boilerplate stuff or or maybe debugging stopping me from running in circles giving me ideas. I I am totally for AI. I'm not on the side of like oh this isn't something that's going to be a thing. It's definitely going to be a thing and if you don't think it's going to be a thing like watch out. But I think these harnesses like like the Cloud Codes, the Hermes agents and and all these things I think these things are going to turn out to be day-to-day staples in in people's workflows. They're just so like productivity enhancers. Like you like from a company perspective if I'm running a company and some like my employees aren't using AI in their day-to-day somehow some way like I don't know how my company is going to survive where everybody else is is using that in some way some shape or form.
Now there's there's things to be said around whether they're security hardened, whether they'll blow up your production database, or yeah production database or a database in general. I mean a lot of the big enterprise companies like the Amazons, the Microsofts they've seen much more outages this year than I remember seeing them in recent years. Whether that's from them having agents kind of do an optimization and saying oh this isn't actually good I'm going to take it down or it's just like the extra AI load that's coming in from everywhere else. Alex Mirran: Yeah, yeah. Narb: No, it makes sense, it makes sense, yeah. It's crazy.
Well I think the other thing too that's crazy is as the hardware gets better the models are getting dramatically better. Like my understanding and like this could be wrong but Mythos right you know Claude's new model they've been talking about and I think even Opus 4.7 if I'm not mistaken were trained on like Blackwell chips. And so this this are these are some of the first models that have been fully trained on Blackwell chips which we weren't having before. Now we have Vera Rubin in production right from NVIDIA. They're producing I mean each chip has 288 gigs of or each GPU has 288 gigs of memory right? And so yeah like you said if if you don't think the models are doing what you need right now I mean just check back a month from now like these things are getting exponentially better. Dario Amodei talks about it the CEO of Anthropic in a really good interview he did with Dwarkesh where he was saying we're in the exponential and that's why you have to be careful about over-investing but we're in the exponential. And so you know as the hardware gets better the software catches up to optimize for what what's available. And I can't imagine I mean some of the the art that some of these frontier labs are understanding how to like post-train how to pre-train how to train these models using more and more hardware. You know it's exciting and exciting to see the the Blackwell chips proliferate and then we're going to start seeing Vera Rubin. I was looking at I think it's like an RTX 5000 it's like 1600 dollars you can get a Blackwell chip at your house like in a consumer GPU. Dell's shipping a a box with that it's like 28 gigs of memory. So yeah I mean that's game changing you know that's that's pretty incredible.
Game changing indeed. And yeah seeing as we're kind of running out of time I just want to get your perspective on a few more things here before we let you go back on your busy day. You mentioned it on the models front, there's been a lot of innovation and these things are getting better and better every week it seems. Both on the closed-source side as well as the open-source side. I guess from from your experience, and I can share mine as well, what what's been your take on like how close open-source and closed-source have been getting and do you think like we'll eventually get to a day where some of these frontier level models the intelligence that they provide you can run on your standard consumer laptop?
Yeah, I think I think we're getting there. I mean I'll I'll preface by saying I'm not an expert in this so you know take my my advice is worth as much as you paid for it here which is nothing. But like Qwen 3.5 35B I think dropped a few weeks ago and it's performing incredibly well on benchmarks and you can run that you know you need you need a chunky amount of GPUs but you could run that on consumer hardware. So yeah we're like slowly getting there. The thing is like the the evals are kind of a black box in a way like everyone has their own eval and then they brag about how good their model did on their eval. You know so it's like we had Swe-bench and everyone was talking about how MiniMax M2.5 and Kim-K 2.5 were performing you know five or ten percent lower than Opus 4.6 for like 90% less cost. And then like Swe-rebench came out or Swe-bench verified I think those are two different ones. And they changed the questions, they changed the bench right and and it didn't perform as well. So I think it's hard to know exactly the performance parity. Definitely like when talking to you know commercial businesses, definitely talking enterprise like they're they're mostly using especially in the west they're mostly using the closed weight frontier models for sure. I mean everyone's using Claude right now it's kind of the thing. That'll probably change but anyway I think what's cool is that we're seeing the open-source labs catch up very quickly. And we actually have some colleagues so a large percentage of the company is in is in Singapore and and different parts of Asia. And so I get we got a lot of interesting perspectives of what's going on and you know especially in the I mean the Chinese labs are very good at using very minimal GPU resources to train models that punch above their weight. They're getting very good at that. And so it's impressive, it's incredible. And so you know I can't imagine that's going to change you know if anything they're going to get better. And then there's some really cool open weight US sort of western models like RC I think it's called Trinity is their their family of models. They're doing pretty well in OpenRouter right now I think. So you know we're seeing the competition heat up in the open weight side. You know on the other hand Llama switched from being open weight to my understanding is they're going the closed weight route right they launched Muse. You know we'll see we'll see. I mean it's kind of interesting like what it means to be open-source. I mean obviously you're giving it to people to use and so but like people aren't necessarily like contributing it like forking it and you know writing PRs and stuff. So it'll be interesting to see how that all kind of evolves. But 100% you can run these models on your computers right now. It's just you're you might sacrifice some performance. But like again you can get a RTX 5000 I think it is with a 28 gig Blackwell chip. I mean you can run some crazy models on that you know and and I think that's like 1600 bucks so it's not cheap but it's like the same as a 4090. Sorry what were you going to say?
Yeah yeah I was just going to say yeah I mean like if well I mean we don't even have the option to buy the the Mac Studio 512 GB anymore right? So and that in itself cost more or less around the same. So I mean yeah I I think I think even so these things might be out of reach for like the everyday person. if you have the money to spend yeah you can set up a really really decent like home rig that kind of future proofs yourself. I don't think these models are going to be taking more resources like they are now. I think people are going to be more clever in the way they they develop these models through like the mixture of agents framework where you only have like a few active billion parameters going at a time even though the model itself is like really big. And like the caching and whatnot. I think all of this is going to kind of come together and make make these models useful and usable versus like a plaything where they they might have been perceived before. Like certainly some like some of the workflows that I have running on on my machines I have a mix of these closed-source and local models where like like you mentioned with UncommonRoute like I I optimize for oh like I I want I want like a really really performant model doing this one thing but I can sacrifice some performance because it's good enough in these other areas. Alex Mirran: Yeah totally. Narb: Yeah you were telling me before the show that you were running some stuff locally what tell tell that again I I would love to hear the sort of setup again.
Yeah yeah for sure. So this podcast in itself if you didn't know is mostly a one-man show so to get some time back in my day to go be able to touch grass I've automated a lot of things post-processing wise. So for example I've been experimenting with kind of different automation flows to produce transcripts and kind of create the clips and whatnot and kind of getting away from some of these paid services that I'm I'm using now. And and like I mentioned it's it's using a mixed source of some some like cloud some Opus but at the same time I've been really jamming hard with Gemma 4 that came out a couple weeks ago, Qwen 3.6 I've been playing around with yesterday. And these things are performing pretty well and like taking like 20 gigabytes more plus minus five of of RAM and just working off my laptop. I'm not even running it on major hardware which is pretty pretty neat to see. I mean over-heat sometimes but I mean that's that's the sacrifice you got to take. Alex Mirran: Yeah, yeah. Your house is probably pretty toasty. Narb: That's right, that's right, yeah. It's still cool where I live so it's it's okay but come summer time maybe I might change things up. But yeah man it's a very very interesting time to be alive like like we mentioned earlier. So I I know I'm really excited to see where we come out the other end of this in like six months' time or years' time from now. And I guess on that point just curious to hear your perspective of where you think Gradient is heading and maybe some alpha on the roadmap that you can share with us today.
Yeah for sure. So we're seeing that like I was saying before we need to add value on top of just inference and training. And so in some cases combining them. So being able to sort of have like a continuously learning agent and then us being able to and then like being able to like very easily deploy that to a GPU through our platform so you can host your own model and you know continuously improve your model. That seems to be where we're headed and that seems to be some of the most valuable apart from routing to the existing models having a model being updated continuously. I think Elon Musk talks about continuous learning. I mean all the labs are sort of thinking about this. And so yeah we're trying to figure out so we're going to sort of take Echo and turn it into a sort of a cloud platform on its own and then take that same tech and use it on Commonstack. And so that's probably the best the best stuff to keep an eye out for. And yeah reach out to me because like we're looking for early teams to sort of test some of this stuff, early partners to like come to us and say hey like I need a model that's good at X, Y, Z thing. Like even even like blockchain use cases right? Like market making or bridging or trading or you know stuff like that like we're interested to hear from you and like know what specifically we think we could solve with post-training because frontier models do a great job.
So yeah that's the main thing. Try you know we have a lot of cool great open-source tech so we're really proud to sort of support the open-source community. You know go try some of that stuff. It's free, use your own, you know the data is all yours like that's a lot of good stuff. And then yeah you know we are we do have roots in the crypto world and we do see a lot of value in crypto-related mechanisms. So yeah keep an eye on our Twitter, keep an eye out, watch our products, watch what's going on and I think you'll see in the long run where we're headed with regard to crypto and on-chain transactions and a token and that kind of stuff.
Excellent. Exciting exciting times. And yeah all that information, all the resources we went over in the podcast today will be in the description below as well as the contact information for Alex and Gradient itself if you all want to reach out. And with that Alex thank you so much man. I really appreciate you taking time out of your busy day to come chat with us today. It's really exciting to learn all about Gradient and catch up and and chat AI. Alex Mirran: No for real thank you so much for having me. I just feel lucky to get on the schedule. You know you have some incredible guests on here so yeah thank you so much. Narb: My pleasure my pleasure man. Yeah anytime. And yeah with that I just want to wish everybody a very happy Friday, happy weekend wherever you may be and we'll catch you back here for another great episode of DevNTell next week. Till then have a good one folks. Cheers. Alex Mirran: Bye.
Listen On
Resources & Links
Share This Episode
Share on XWatch Episodes Live!
Subscribe to our event calendar and never miss a live episode.
View Event Calendar