The Future of Software Engineering: The World After AI
A staff engineer on where AI in software is now: context limits, typesafe models, decompiling binaries, and juniors who won't be mentored the same way.
I became a software engineer in 2017. I remember getting my first high paying job at a tiny startup in Los Angeles, being green and not knowing much about what it meant to be a software engineer. As the years passed I became more and more valuable to companies: I became a senior software engineer, then a tech lead, and then finally a staff engineer. Right around that time, ChatGPT 3.5 was releasing and people were going crazy about it.
How I got here
The first time I used ChatGPT I was impressed. I remember comparing it to my past experiences with the old chatbot site, where the conversations would lose track of context mid sentence and most of it was unreadable. Then suddenly, you have this machine that can not only string paragraphs together, but have a conversation with you.
AI was always a trend that I was tracking as a fan of sci-fi. Then it became real, and everyone was wondering how to use it. And as always, people rush in and use it the wrong way.
I saw people putting all their faith in it, thinking it would do the job of multiple people at once. I had a manager that was pushing hard for that kind of approach, and his tool for triaging issues was always hallucinating logs, entities and services that didn’t exist.
It is just a tool, not a replacement of thought and judgement. And that is where I think we must focus as software engineers: thinking things through, planning, deliberate design, avoiding accidental architecture, and also knowing when to trade control for speed. The engineers that can do that will dominate.
That gap shows up in the data too. In the 2025 Stack Overflow Developer Survey, 84% of developers said they use or plan to use AI tools, but only 32.7% trust the accuracy of what they produce. The top frustration, for 66%, is answers that are “almost right, but not quite.”
- use or plan to use
- 84%
- trust its accuracy
- 32.7%
source: Stack Overflow Developer Survey 2025
What is happening now with AI and software?
As a software engineer, I am tracking a few things:
- New models coming out, and their ability to code better and better with each iteration.
- Different kinds of typesafe models like Jev, which are cheap and fast and can be used for routing decisions in code.
- One of the more interesting and, in my opinion, worrying trends: AI being able to reverse engineer everything by decompiling binaries.
Let’s go over each one at a time.
New models: the context is still the limit
New models like (at this time) Claude Opus 5.5 and GPT Astra are very impressive. However, the context is still the limit, and I don’t mean just the context window, but how the context gets fed to them, how they access information.
With things like MCP (Model Context Protocol, “an open-source standard for connecting AI applications to external systems”) providing connections to external tools, these models have become much better and more useful. Being able to have the LLM query the DB and check for bugs in endpoints while checking monitoring tools is an amazing time saver in triage and debugging.
The code quality is amazing. Most of the time it is better than what I would write, and I think for most basic apps that vibe coders like to build, it is sufficient. LLMs are good at individual tasks, but when you get them working across systems, they suck.
The gap here is that the LLM will not understand the business. You can give it access to Confluence or whatever knowledge base your business uses, and it can come up with some excellent suggestions for certain features and improvements to the system. But it won’t ever know everything, because not everything was documented.
So the more chaotic businesses have a disadvantage when it comes to harnessing the power of LLMs to their fullest, just because there wasn’t enough effort put into documentation and knowledge transfer, the same thing that makes onboarding hectic. The more chaotic the company, the safer you are, as long as you have become a valuable member of the team and hold knowledge that the LLM wouldn’t be able to easily find or investigate, which is the case in most distributed systems.
I had to work on a system that had a dependency on a very old CMS that was being misused for feature flags. That code had three different ways of getting feature flags, and none of them was clearly labeled “feature flag”. The LLM just couldn’t find the thread and the right place to look.
“From what I’ve seen they’re great at identifying trees and bad at mapping the forest. … The real challenges are things like identifying and mapping out any instances of temporal coupling, understanding implicit business rules, and inferring undocumented contracts and invariants. And LLM coding assistants are still pretty shit at those tasks.”
Small, specialized models are the trend to watch
On the different kinds of models coming out, I think this is the trend to watch. LLMs are great, but models that have more specific uses and can become part of software composition (they need to be cheap and fast), like Jev, can open the doors to solutions that LLMs are not good or efficient at.
Jev is a typesafe model made by TypeSafe AI. It doesn’t generate text: you send it a state and typed questions, and it returns typed answers with probabilities. Per its docs, input costs $0.042 per million tokens and output tokens are free.
For example, I use Jev on IndieDex.gg to help decompose users’ queries when my initial regex lexical filters fail to break them down. Before, I was using DeepSeek V4.1 Flash, and it was doing a good job, but it was too expensive and too slow. Once I changed it to Jev, my costs got 6x cheaper and it got 10x faster. Here is the benchmark I ran:
| Metric | DeepSeek | Jev |
|---|---|---|
| Extract p50 | ~1.5s | ~130ms |
| Full route p50 | 2871ms | 1256ms |
| Weird title-resolve | 96-100% | 100% |
| Hard-filter agree | n/a | 100% |
| Weird top-N overlap | n/a | 100% |
| Cost (fixture) | ~$0.014 | ~$0.0025 |
extract p50
- deepseek
- ~1.5s
- jev
- ~130ms
full route p50
- deepseek
- 2,871ms
- jev
- 1,256ms
cost per fixture run
- deepseek
- ~$0.014
- jev
- ~$0.0025
source: my benchmark on IndieDex.gg
So extract got roughly 10× faster, the end-to-end route roughly halved, cost dropped a lot (around 6x), and quality held on my fixture set.
LLMs are too big and too expensive
In general, I believe other kinds of models are what will turn the world upside down. LLMs are great, but they are too big and too expensive: see how their subscriptions are subsidized. The numbers in that SemiAnalysis piece: despite being about 10% of Anthropic’s revenue, subscriptions “can take up over 40% of inference compute,” and their model puts a maxed-out Opus 5.5 subscription at a -369% gross margin. It is not sustainable for that to continue.
And have you noticed that whenever a new model is coming out, or has just come out, the models you used to daily drive suddenly get dumber and slower? That is because they need to shift the compute around to train and maintain the new models, and the resources for GPUs and memory are very limited.
The labs deny the “dumber model” part. After users reported Claude getting worse in 2025, Anthropic published a postmortem blaming three infrastructure bugs and stating: “We never reduce model quality due to demand, time of day, or server load.” A 2026 follow-up traced another round of complaints to Claude Code changes, not the models.
I think they are BSing. It makes sense that they would “shrinkflate” the model whenever possible.
So models like Jev, that have a clear purpose and are integrated into the product, make more sense from a financial perspective. They are a great way to bring intelligence into software without needing something that generates so many tokens. It responds with a typesafe response: it either says “yes or no”, or it scores, or it chooses from multiple choices.
And even in other applications besides software, like health care imaging, video and image generation, and music generation, all these specific models are making leaps of their own. If you use Instagram, you must have seen a lot of parodies done with AI lately, with Seedance and MiniMax H3. The leaps in quality are astounding, and they have also become cheaper to run: MiniMax H3 can run on consumer hardware. MiniMax H3 is an open-weight model, and with ComfyUI’s day-0 support its memory needs dropped from 123.6 GB to 42.5 GB, enough to run locally on a GPU like an RTX 3060 with offloading.
This is the opposite trend of LLMs, which seem to become more and more expensive. And at some point, the subsidizing will need to stop. When that happens, what will become of the businesses that now over rely on these technologies?
I don’t believe this will happen overnight. I believe it will start with models becoming slower and dumber to offset the cost of subsidizing them, much like our products get shrinkflated over time to account for the dilution of value of our currency. I believe smaller and specific models will become the tools that stay accessible and get more and more integrated into our workflows.
So not an AGI or ASI taking over your job, more of a change in tooling and skillsets and design patterns.
Decompiling binaries: the scariest one
Now on the last point, and the scariest of them all: the capability to decompile binaries. Decompiling means turning a compiled program (the binary you actually ship) back into readable source code. It used to take a skilled reverse engineer weeks; tools like GhidraMCP now let an LLM drive a decompiler directly.
A lot of people are having fun with it, mixing and matching games, like putting Minecraft on Elden Ring and so many other examples. It is like modding on steroids.
However, that is not all it can do. Now all software can be cracked, analysed, rebuilt and breached. This means the delivery of apps will likely shift more and more towards web platforms and servers. We have already been seeing that, but now it is accelerated. There is no point releasing a desktop app with compiled binaries that can be reverse engineered: that basically means you are shipping out the source code every time. It increases the rate of zero day vulnerabilities being found, of copycats figuring out some of the secret ways patented algorithms work, and piracy in general.
This is already happening at scale. In April 2026, Anthropic reported that its Claude Mythos Preview model can take “a closed-source, stripped binary” and reconstruct plausible source code for it, and that it had found “thousands” of high- and critical-severity vulnerabilities, with “fewer than 1%” fully patched at the time (Anthropic).
“The juxtaposition of Claude+Ghidra being able to take apart understand and reimplement the core features of this thing in hours while also having to babysit it “no, those encrypted packets going over the CAN bus aren’t from wifi” and “please actually look at the code you just decompiled instead of guessing how they work” was pretty amusing.”
— colechristensen on Hacker News, on reverse engineering a Tesla Model 3 computer
Something I am watching closely is the AnyPS5 project, which is reverse engineering the game binaries and Sony’s system libraries: basically game code repackaged, Sony stuff replaced, GPU stuff translated. It is still far from delivering a full PS5 emulator, but I believe it will get there, and more advancements in LLMs, and maybe some specialized models, could accelerate it.
This could mean that all hardware could get cracked and jailbroken, and that could plummet the revenue for many companies. A great day for pirates, a bad day if the way your salary gets paid is as part of that revenue. Do you really think the execs would take a pay cut?
Protecting IP when everything can be decompiled
I think server side as much as possible, maybe some new form of code with encryption if that can be done performantly, but it does feel unprotectable currently.
For indiedex.gg, it is all served, so I’m not worried. I have also written about how it works, and those articles are all out there somewhere. I think the workflow of how I classify/tag data and synthesize things is the most unique thing about it, and to reverse engineer that you need to access the database and analyze it. It is not baked into a binary.
The security risk
It scares me more as a user: there are so many dependencies we have no control of. Teams shipping software will have to be held to much higher standards of security, and their processes and operations will need to evolve to include LLMs trying to crack the code. This could become a whole new sector of cybersec.
Where engineers are going: knowing where quality matters
I spend a lot more time thinking about where quality matters and where I am willing to accept some slop.
In the UI, for example, it has become super cheap to just strip out all components and switch design systems. As long as there is robust testing, you can have the LLM do full coverage of tests, and with the right QA process you can trust those issues would be caught before they hit production.
But on critical paths like payment systems, you do want to triple check things and ensure everything is hardened and covered and the tests make sense. You need to put a lot more effort into reviewing and supervising the critical paths, to be methodical about what matters to the business.
A broken button that paginates an image? Acceptable. A duplicate charge on a credit card of a client? Not acceptable. And of course, for important things like the code of micro computers that read sensors in aircrafts and software for pacemakers, you don’t want to vibe code that.
| Where | Is some slop OK? | What it takes |
|---|---|---|
| UI, design systems | Yes, if tests and QA catch it | Robust testing, full test coverage by the LLM, the right QA process |
| Payment systems | No | Triple check, harden, review and supervise, make sure the tests make sense |
| Aircraft sensors, pacemakers | Never | Don’t vibe code it |
Some engineers think the role itself goes away:
“The future of engineering is product management. I don’t believe there is any world left for people whose primary responsibility is opening pull requests”
I partly agree. I think it will be harder to spot the lazy engineers that over rely on AI, but for them, it means AI has already replaced them, and eventually it will catch up to them.
The entry-level problem: less mentoring, fewer jobs
We have a gap in skill forming: entry level engineers are not getting hired, and this will lead to an eventual gap in the workforce. The jobs data is already tilting. Per Indeed Hiring Lab, only 4.5% of software development job postings were entry-level in Q1 2026, while 69.3% were senior.
- senior
- 69.3%
- entry-level
- 4.5%
source: Indeed Hiring Lab, July 2026
And in LeadDev’s AI Impact Report 2025, 38% of engineering leaders agreed that “AI tools have reduced the amount of direct mentoring junior engineers receive from senior engineers.”
I had coworkers that were obsessed with Uncle Bob and SOLID principles, and they taught me how to write better code through reviews and pairing. That is something I don’t see happening as much now that I work in a very chaotic environment with a “just get it out fast so we can move on” approach.
Microsoft’s Mark Russinovich and Scott Hanselman make the same case in Communications of the ACM: AI gives senior engineers a big boost while putting an “AI drag” on early-career developers, so companies “hire seniors and automate juniors”, and the pipeline that produces the next seniors “quietly collapses” (InfoQ summary). Their fix is a preceptor program pairing juniors with experienced mentors on real teams.
“I’ve already seen an instance of a junior developer being completely unable to fix vibe coded technical debt, even with guidance, because AI analysis is fundamental to how they understand code”
Not everyone sees it that way. Francisco Trindade, a VP of Engineering at Braze, wrote that an intern led the development of a whole feature on his team: “Training actually became much cheaper.” His take: “AI expands what every level can handle, including juniors.”
The first story is closer to what I see. I think most people in general avoid thinking when they can, and AI enables that kind of behavior. I have been guilty of that. It is hard to find balance when you are stressed out and there is this thing that you can just prompt your problems away at. It is always the path of least resistance.
What makes good code good
I think the biggest thing is understanding how to decompose, and how to use the small pieces to compose it differently. The biggest principle behind that is the single responsibility principle. If I had to teach one thing to an entry level engineer, it would be this. It makes the code more readable, easier to test and easier to reuse.
In my experience, AI can create a lot of repetition if you leave it unchecked. I do have rules about single responsibility for my agents, and about documenting as you go, so it happens a lot less with the right setups and audits. But it can still happen: the AI will create some redundant functions that can later become pesky bugs, because it did the same thing in multiple different ways and to different standards.
So having your preferences, principles and standards locked in matters a lot if you are going free form with the code.
What I’d tell someone starting out
Learn to think as an engineer, even without the software aspect. Designing systems that solve problems is a very transferable skill. Be agnostic about your methods: it can be business, marketing, all of it applies.
I also think it is always a good idea to diversify your skills. Work on some evergreen skills like marketing, conversion optimization and lead generation: things that businesses will always need, and skills that are enhanced by your knowledge of software.
And use your software knowledge as an advantage over the average vibe coder that never had to center a div.
Software engineering isn’t going anywhere
Software has changed forever. But software engineering isn’t going anywhere as long as software exists. Tools may change, demand may change, but talented thinkers and problem solvers can translate their talents anywhere.