★ FEATUREDBuild your BI dashboard from scratch — faster than taming Power BIRead →
Platform

The AI built it in an afternoon. Three weeks later nobody knows why.

There’s a word for it: vibe coding. You describe what you want, the AI writes the code, and you barely glance at it. It works — and it works surprisingly well.

84% of developers do it now. Only 29% trust the result.

That number is the whole story.

Why everyone does it

Because it delivers. Tasks get solved 20-45% faster. New developers complete a project up to 55% faster than the old way. Four out of ten new SaaS products this year are built this way.

We do it ourselves. Every day. This isn’t an article against vibe coding.

What happens three weeks later

Something needs to change. No one remembers why it was built that way. The decision was made in a conversation that’s been scrolled away, and the agent who made it no longer exists.

It’s not a feeling. It’s measured:

  • Technical debt increases 30-41% after a team starts using AI tools — measured across 8.1 million code changes.
  • 45% of AI-written code contains a vulnerability from the industry’s top-10 list. Tested across over 100 language models.
  • Secret keys leak twice as often as in hand-written code.
  • 20-30% of a team’s time is spent after three months fixing errors from the AI.

And the worst part: the debt is invisible. The code passes tests. It gets through review. The problems lie in error handling, edge cases, and security boundaries — places that only reveal themselves under real load, with real users.

It’s not the AI. It’s the lack of a harness.

A horse without a harness isn’t a bad animal. It’s just not harnessed for anything.

Vibe coding doesn’t fail because the model is bad. It fails because there’s nothing else: no plan that survives the conversation, no decision that’s remembered, and no one saying no before something goes live.

That’s what we’ve built.

The plan is written before the code

Every feature gets a written document before the first line of code: what it should do, what it must not do, and how to tell if it works.

An agent that gets a plan builds what’s written. An agent that gets a conversation builds its own interpretation — and you only notice the difference when it’s coded. The plan is also the only thing that still exists three months later when someone asks why.

Decisions aren’t reopened

Every team has choices that were made. We don’t use that framework. We don’t touch payments without a direct order. Customer data never leaves the EU.

Making them isn’t the problem. Forgetting them is — and then the next agent suggests something already rejected, with a good reason no one remembers.

With us, they’re stored in one place and delivered to every agent when it starts. Along with what disappears first: what was chosen against, and why. That’s where the wasted time lies.

"Done" requires proof

An agent must not declare its own work done.

When a task is marked complete, a real browser opens the page and checks: is the button there? Does it work? Does the screen look right on a phone? Not "the tests are green" — but proof that it works.

If it fails, the task goes back. It never reached your users.

That’s exactly what the 45% is about: code that passes tests but still carries a vulnerability. A test that only asks "did it answer?" can’t see it. One that looks, can.

Your agents. Your machines.

Agents run where you want them: on your own computers, or in an EU cloud in Stockholm. Your code never leaves your control, and no third party gets a peek.

For some, it’s a preference. For others — healthcare, finance, the public sector — it determines whether the project can even start.

And you control it from your phone

You don’t code from your phone. You direct.

Start an agent anywhere, watch it work, send it a message, see what it decides. Get a vibration when it’s done — or when it needs a yes from you.

Built at 2 AM. You nod from your bed.

When vibe coding is right — and when it’s not

Honestly, with the numbers in hand:

Use it freely for a prototype, an internal tool, an idea to test this week. The speed is real, and the debt doesn’t have time to cost anything before you throw the code away anyway.

Don’t use it without a harness for anything that needs to last: payments, customer data, a platform you’ll maintain in two years. Here, the 30-41% isn’t a statistic — it’s your calendar next spring.

Also note who benefits most: experienced developers report 81% higher productivity. Juniors show no improvement at all because they lack the judgment to evaluate what the AI delivers. A harness is what gives an entire team that judgment — not just those who had it to begin with.

500 hours — designed while it was used

Cardmem didn’t emerge by accident. It was designed, through roughly 500 hours of work side by side with the agents it controls.

But it was designed in one specific place: right in the middle of real customer projects. We’re builders — we build commercial, customer-facing platforms that need to last, and we build them with the exact engine we sell. Every rule in the harness was drawn while a real deadline was running, and tested on work someone paid for.

That’s the difference between a method and a theory. The rule that the plan is written before the code wasn’t found on a whiteboard — it was drawn, used, found insufficient, and tightened because an agent had built its own interpretation of a conversation and we only discovered it once it was coded. The requirement for a real screenshot before anything can go live came from tests being green while the page was wrong. The requirement to prove that a control was even run came from a control that returned the same answer no matter what — and that we trusted.

We take our own medicine every single day. If something goes wrong for us, it becomes a blocker — and that blocker comes with you. That’s why it works in practice and not just on paper.

It’s not a theory. We build this way ourselves.

FD Sundhed is a full healthcare platform — booking, payments, customer portal, employee access — plus native apps for both iPhone and Android.

Eight weeks from first line of code to real users. One person was behind it, and he doesn’t write code himself: he directs 15+ AI agents through the engine described above.

Want to see it?

We’ll set it up on your own project, not a demo. That way, you can see your own code go through the loop and judge for yourself if this is how you want to work.

Read how the eight weeks went.

Hi — I'm Aidan