18 KiB
AI Coding Rate Limits are RIDICULOUS Now - Here's How You Keep Scaling Anyway
Sursa: https://youtu.be/NZq88JAJSag?si=hx_PHjK5_YDxxFR_ Data: 2026-09-24 Creator: Cole Medin Format: Video (~14:08 min) Limbă: en Tags: @coaching
TL;DR
Rate limit-urile pe Claude Code/Codex s-au înăsprit tot anul; Cole Medin (canal AI coding) a testat zeci de combinații de modele pe workflow-uri complete (plan → implementare → review) și a găsit un pattern clar: treapta de plan/review beneficiază de cel mai capabil model (Claude fable 5.1 / GPT-6 astra), dar implementarea — partea cea mai consumatoare de tokeni — merge la fel de bine cu un model mic/ieftin (GLM 5.3 flash), cu ~4x mai puțini tokeni și rezultate comparabile sau mai bune. Recomandarea lui: model puternic doar la plan+review, model ieftin la implementare, indiferent de unealtă.
Linkuri: https://youtu.be/NZq88JAJSag
Puncte cheie
- Cel mai important pas al unui workflow AI coding e planul — un plan bun face ca implementatorul să nu mai aibă nevoie de un model de top.
- Implementarea consumă cei mai mulți tokeni → acolo se face economia reală de rate limit, nu la plan/review.
- Testul lui: același joc construit cu (a) doar modele open, (b) doar Claude, (c) model puternic la plan/review + GLM 5.3 flash la implementare — varianta (c) a ieșit la fel de bună sau mai bună, cu ~4x mai puțin cost/tokeni.
- Pentru task-uri mici, bine delimitate, cu om în buclă → poate merge un model ieftin pe tot fluxul.
- Pentru task-uri foarte dificile/neexplorate → merită model puternic peste tot.
Insights
- Aplicabil la fluxul Ralph (proiecte autonome, planning → PRD → implementare per story): planning/PRD deja pe Opus, implementarea pe Sonnet — pattern-ul descris confirmă alegerea curentă. De văzut dacă la un moment dat merită încercat un model mai ieftin/open doar pentru story-urile de implementare simple, ca să reducă presiunea pe rate limit. @work @project
Transcrierea
This week I've hit my Claude code rate limit way faster than I ever have before and I'm not even doing that much more work. This is my max $200 a month plan and the worst part is I also have a $200 a month pro subscription for Codex and I've almost exhausted that limit as well and both of them aren't going to reset for another three or four days. And so if you thought it was just you hitting your rate limits faster don't worry it's not. We're all suffering from this and unfortunately we knew this was coming it's been a trend this entire year our rate limits getting worse but it's just finally come to a point where I can't just be more efficient with my token usage. I mean it's ridiculous that I have to wait three to four more days before I even have a reset on my max plans. So over the past half year I've been doing a lot of experimentation with using different models especially open ones in my AI coding workflows but now that's gone from just experimentation to okay I absolutely need this as a crucial part of everything that I'm doing. I mean look the proof is right here and that's not just me as an industry when we get more into harness engineering and loop engineering and software factories whatever your AI coding workflow looks like we're trying to scale the output with our coding agents but now even more than before the biggest limit is just the number of tokens that we can spend. That is why we need to strategize and figure out for our larger AI coding workflows not single agent where do we absolutely need the largest and most capable LLM versus other places where we can use something that is smaller cheaper faster and still get basically the same output quality and yes that is possible. Trust me I've been deep in the lab this week expanding on all the research I've been doing this year figuring out for a traditional AI coding workflow where you have your planning implementing validating and reviewing where do we need the best LLM versus where do we not that's the big question to answer right now so you don't hit your rate limit three days into your weekly reset. So for all the testing that I've done this week I've done it with my AI software factory I've covered this system a lot on my channel recently it takes all my best practices for AI coding with planning and implementing and validating it packages it up into a complete system where I can send in a scope of work and it just builds and validates everything without me in the loop at all it's the ultimate test of how autonomous can we make AI coding while still keeping it reliable but more importantly for this video it allows me to very easily test different combinations of large language models with the full development cycle so what I've been doing is building a bunch of different applications with the AI software factory including this video game that I'll show you today just makes for a nice and visual core example but all the other things that I built the results I got is pretty much the same as what I got with the different combinations for this game so building the same game but with only open models using Claude and using Codex a lot of other testing that I did as well but these are the three core columns and so essentially I just took the best model from each group right with open models deep seek v4.1 flash is pretty much the best of the best so if we look at like the live llm benchmarks here as far as open models go with that open tag deep seek v4.1 flash is the best right now and then for Claude Claude fable 5.1 max effort is the best and then for Codex GPT6 astra max effort is the best and so going back here for my core testing with open models I'm using deep seek flash for planning and reviewing and then glm 5.3 flash for the implementation and I actually did this as a baseline where I used it for implementation for every single test and then for Claude using fable for planning and reviewing and then for Codex using astra the most important thing here is you'll notice I'm always using the more capable llm for planning and reviewing and then the smaller model for building this is something that I've determined in my experimentation this week and a lot of what I was doing earlier this year as well the most important step of your ai coding workflow is always the plan step if you have a well written plan the implementer following that it doesn't even have to be that capable of an llm to pretty much get the best results possible right like using fable for planning and implementation it's you essentially the same results as using fable for planning and sonnet for implementation so that you'll see that very intentionally laid out in all the testing I show you today and by the way for implementation generally that's the most token heavy part of your workflow so for using the smaller model for that that is what prevents us from hitting our rate limits super quickly now obviously for some work if it's very very difficult and unexplored you might want to use a large model for everything and I don't do this in my software factory but if you do have human in the loop really small and bounded tasks you can use an open model for literally everything so there's different combinations that make sense in different situations but this is generally what I recommend and I'll show you a lot of my testing that has led to this but for me my favorite flow right now my favorite combination is to use GPT6 Astra for the planning and the review and then GLM 5.3 flash for the super fast and cheap implementation this is my core workflow for everything that I'm doing right now with my software factory and even outside of that the sponsor of today's video is scrimba and they just released something incredible called explain you ask a question and it gives back a fully narrated video lesson in two to three seconds built live while you watch and it plugs right into codex and claud code so now over in claud code I do slash mcp we can see I have scrimba explain connected it was just a single command to do this and so now I can ask claud code to make an explainer for anything and so I thought this would be a neat demo here to show it making a scrimba explainer for a more complicated pull request from my open source project archon so here claud reads the diff it builds up the lesson and then it streams the scrimba slide by slide so I get the link right away I can watch it as it builds the explainer and I gotta say the explainer that I got back here at the end it genuinely impressed me let me show you so this is the original pull request that I made the explainer for and then here is the video so I'll just shut up for a second and play a couple of the clips here PR 3416 gives three adapters one shared way I'll go to a more high level diagram here is the contract that the cli adapter and the web adapter implement the work goes into the code as well without an error click a little bit further here lives in core the helper accepts any object which lets the headless platform pass its structurally identical work so I'm not going to play the entire thing obviously but yeah it sounds great and it's breaking things down really nice and simple for me I love it explain also works from chat gbt and as a chrome extension for any article give explain a try for free with my link in the description now the question you might have at this point is how exactly do we build a larger ai coding workflow where we're using different models and providers because maybe you want to implement the same way I am you want to experiment like I am whatever it is there are a lot of different ways you can do this the most manual but simple way is just to have each coding agent output a handoff document as markdown so you feed that into the next model or the next coding agent and so you just go through one at a time opening up different coding agent sessions there are also harnesses out there like omniscient that make it very easy to work with models and different providers in this way but what I do for my ai software factory and my coding workflows in general is I simply use my open-source harness builder arcon this makes it really easy for us to build these workflows that I've been running hundreds of times the last week to combine different models and providers so for every single test that I ran every single combination of models the shape always looks like this as an arcon workflow and I'll show you what the workflows look like in a little bit so the input is always a github issue that describes what we want to build then I have the most powerful model do the planning like GPT-6 astra as my favorite we take that plan and then within the arcon workflow we automatically pass that into the next node which is glm 5.3 flash doing the implementation and then we have astra review and then any findings that sends back to glm 2 correct and we have that loop with obviously max retries and then we have the check at the end which is just running our tests making sure the build is green before we do the merge at the end of the software factory and so this is a pretty traditional AI coding workflows nothing crazy here this is the exact shape for every single test I ran there are also a lot of different ways to access open models for our AI coding workflows but my favorite recently especially because arcon supports this is to use pi as the AI coding agent harness and then for accessing the LLMs using neon's AI gateway I've always loved neon I love their postgres database solution they've been adding a lot of other really cool AI features as well like AI gateway and all their other new backend features like object storage authentication and functions but yeah for right now I'm using AI gateway which gives us really reliable and at cost access to pretty much every open model that you could hope for so for example we have kimi k3 uh what else do we have here we got glm 5.3 flash which I've been talking about a lot and so this is just my easy way to access all the models that I need within my arcon workflows and with pi but regardless of how I am accessing these open models the important thing is that I am I'm not just using Astra or fable for everything like a lot of us are tempted to do and you can do the setup with anything you don't even need arcon it's my tool it's free and open source and I use it to run the AI software factory but also I'm not telling you you have to go use my tool at all right like this is just what I've been doing for my experimentation to give these results to you and help you think about what models to use what combination and so for my arcon workflows this is a bit of a simplified view but pretty much every single one of them looks like this right we have the planning step at first where we're using the more powerful model and then we take that plan markdown document as a handoff to go into the build step with the cheaper model like glm 5.3 flash through neon and then review with more powerful again and then fix anything that came up with the cheaper model this is the shape even beyond my experimentation here for just generally how I work with coding agents now okay so I wanted to spend good time explaining my workflow my experimentation and the models that I recommend but now I want to get into the actual game with you as the core example here just to give you a visualization of the results that I've been getting across all the different apps I've been building this again like I said at the start of the video is very much the results I got no matter what I was building and so the first version of the game looks terrible you can probably guess this is what I built with only open models and so by the way this game is an idea from a friend it's really cool so I thought I'd put this through the factory it's a weather chasing like storm chasing game on Neptune it's actually multiplayer we can have different people join and man the stations here I won't set that up right now but it's pretty comprehensive game that it gave us that prd into the software factory and so this version of the game I used open models so deep seek v4.1 flash for the planning and reviewing and then for the building I use glm 5.3 flash now the visuals are so bad you can't even really tell that I'm moving there's a couple of clouds you can see outside the windows but yeah you can see the heading changing in the top left as I'm turning around as the pilot but overall this isn't even really a good starting point for a game and then for the next version of the game this one was built entirely with Claude because I wanted to show you an example here where we aren't using open models at all I used Claude fable 5.1 to build this entire thing and yes for the sake of tokens I couldn't do something super comprehensive so I unfortunately wasn't able to make something that looks like super beautiful but I mean this still looks pretty good definitely a lot better than just the open models and let me go over to the pilot here I wish I could run but I'll go over to the pilot and show you that things look a lot better like if I take the controls here and I change the heading you'll see that like I mean it's a little bit choppy but I'm moving towards the clouds and then when I go forward the clouds are actually getting closer like the navigation on the planet is actually working pretty well here it'd be cool to show you all the other stations as well but it would take a lot of time the point is that Claude fable 5.1 it did a pretty good job for the very initial proof of concept for this game I didn't allow that many tokens but yeah it's pretty good the next version of the game I am genuinely impressed not that it's like actually amazing but the fact is I used glm 5.3 flash to write every single line of code in this version of the game and then for the planning and reviewing I used GPT-6 astra and so this version of the game was about four times less tokens or like less cost overall because I'm using glm for the implementation and it plays pretty much the same actually the moving is even smoother here going towards the clouds I would say this version of the game actually is better even though I was using an open model to write every single line of code so of course it helps to have astra review but I'm still saving a ton on my cost slash rate limit by using this combination of models and so yeah a little bit of a cheesy game very much just a proof of concept but the point here is this is the results that I got for pretty much every application that I built and remember going back to the live bench here Claude fable 5.1 and GPT-6 astra they're pretty close in capability and so it's pretty awesome that we were able to build even a little bit of a better game and other apps with this combination compared to just using Claude that is the proof you need you don't always have to have the best large language model for every step of your process it is well worth your time and I hope I've really proved it here to analyze your AI coding workflow figure out where you need the best model and where you don't because that is what's going to prevent you from jacking up your rate limits for your subscriptions like Claude and Codex and so my AI software factory is how I've been doing all my testing how I run my workflows now but whatever you're currently doing this this lesson applies like you should really explore open models at this point because it's becoming clear that we can't always run on our frontier models so I hope that you found this testing interesting as I'm going through all these different combinations building out different apps if you did I would really appreciate a like and a subscribe follow along as I continue to build out my software factory and test these combinations and with that I will see you in the next video