Overslaan naar content

AI Engineer - iO Bonzai

  • On-site, Remote, Hybrid
    • Amsterdam, Noord-Holland, Netherlands
    • Antwerp, Vlaams Gewest, Belgium
    • Brussels, Brussels, Belgium
    • Copenhagen, Hovedstaden, Denmark
    • Den Bosch, Noord-Brabant, Netherlands
    • Eindhoven, Noord-Brabant, Netherlands
    • Ghent, Vlaams Gewest, Belgium
    • Göteborg, Västra Götalands län, Sweden
    • Herentals, Vlaams Gewest, Belgium
    • Rotterdam, Zuid-Holland, Netherlands
    • Sofia, Sofia, Bulgaria
    • Stockholm, Stockholms län, Sweden
    • Utrecht, Utrecht, Netherlands
    +12 more
  • Technology

Job description

iO's ambition is to become the AI Native Marketing & Technology partner of the next decade. iO Bonzai is the platform that has to carry it, and right now it still answers questions: one question, one answer, one model. We are moving it from chat to work, and from work to AI team members that pick up tasks on their own.

That shift is this role. You will first work on the new work harness: it takes an instruction, breaks it into subtasks, hands those to specialised subagents, tracks progress and returns a finished result. Then on autonomy — where agents stop waiting to be asked and start acting as team members with a responsibility of their own.

You build the platform, not the use cases running on it. Those come from 2,000+ colleagues inside iO and from the organisations that buy iO Bonzai. The team is small and remote-first, spread across iO's campuses in the Netherlands, Belgium, Sweden, Denmark and Bulgaria.

About iO Bonzai
iO Bonzai is our own AI and automation platform: Chat for secure assistants, Apps so teams build their own tools, Connect for access to CRM, ERP and CMS. EU hosted and built for the EU AI Act. The full architecture is a mix of components including open-source forks, hosted primarily on Azure. The software is primarily TypeScript, Python and Rust, gradually being replaced with our own technology.

What you will do

  • You build the agentic harness: task decomposition, subagent orchestration, task planning and progress monitoring, plus the context and memory management that keeps a long run coherent. Then you make agents autonomous: checkpointing and resumability for long runs, agents triggered by an event rather than a person, and approval gates for anything with external impact.

  • You work across the platform's core components rather than inside one of them:

    • The API gateway:model access, routing, failover, usage and cost control. You make the trade-offs on model routing: which model serves which task, how to balance cost against capability, and when to cache, fall back or reject.

    • The agent runtime: orchestration, subagents, task planning, long-running execution.

    • The Workflows platform (based on n8n): predictable and repeatable workflows connecting actions across platforms and LLMs.

    • The MCP hub and tool layer: how agents reach systems, and the permissions governing it.

    • The sandbox: running agent-generated code without creating a liability.

    • The retrieval and knowledge layer: embeddings, vector search, organisational memory.

    • The observability and evaluation platform: tracing runs, and measuring whether they did the right thing. You define what "correct" means for agent and workflow output. You design evaluation criteria and benchmarks that measure quality, consistency and safety across run, not just the tooling that runs them.

    • The guardrail and policy layer, keeping personal data and credentials out, anomaly detection, audit trail. You design the governance model for the platform: which guardrails apply to which use cases, how EU AI Act requirements translate to technical controls, and where human approval gates are needed.

  • Frameworks for the agent layer are still open, and that decision is partly yours.

Questions? Our recruiter, Fleur, will be happy to help you. You can apply until Friday, October 2.

Job requirements

  • A relevant IT education at bachelor level or above in AI or software engineering.

  • Deep familiarity with LLM architectures: you can explain how they work and you keep testing assumptions rather than treating them as settled.

  • A strong software engineering foundation, with the architecture skills to design a component end to end and defend the trade-offs, we work in TypeScript and Python, and do not expect fluency in both.

  • Experience building the harnesses and operating systems that produce AI output: orchestration, tool use, multi-step and long-running execution. You have worked on guardrails or AI safety in production.

  • You have an opinion on what good AI output looks like and how to measure it.

  • You think about cost, compliance and risk as engineering constraints, not afterthoughts.

  • You can defend an architecture decision to a room that includes engineers, product owners and compliance.

  • The ability to balance feature delivery with foundational work on maintainability, stability and security.

  • You work AI native, coding agents are part of how you build: Claude Code, Cursor, Copilot or their successors.

  • Fluent English, the working language of the team. Where you live is up to you.

Nice to have

  • Built MCP servers rather than only consumed them, or run untrusted code safely in sandboxes.

  • Built evaluation and benchmarking for AI systems, measuring correctness and consistency across runs rather than eyeballing output.

  • Azure experience, and worked with tracing or observability for LLM systems.

On-site, Remote, Hybrid
  • Amsterdam, Noord-Holland, Netherlands
  • Antwerp, Vlaams Gewest, Belgium
  • Brussels, Brussels, Belgium
  • Copenhagen, Hovedstaden, Denmark
  • Den Bosch, Noord-Brabant, Netherlands
  • Eindhoven, Noord-Brabant, Netherlands
  • Ghent, Vlaams Gewest, Belgium
  • Göteborg, Västra Götalands län, Sweden
  • Herentals, Vlaams Gewest, Belgium
  • Rotterdam, Zuid-Holland, Netherlands
  • Sofia, Sofia, Bulgaria
  • Stockholm, Stockholms län, Sweden
  • Utrecht, Utrecht, Netherlands
+12 more
Technology

or