Skip to content
nothiring. blog What an AI Engineer actually does hire with us hire

What an AI Engineer actually does

What the role does, what skills it asks for, and why one title covers four dimensions that look nothing like each other.

You hear the AI Engineer title everywhere, and there is more to it than it looks. If you are hiring one, or you just want to understand what the job is, read on.

Six LinkedIn job ads with AI Engineer in the title: Forward Deployed AI Engineer at Quantrue, AI Engineer for internal automation at Fever, Applied AI Engineer at Dwelly, AI Engineer for software delivery enablement at Getnet Platforms, AI Native Product Engineer & Orchestrator at Accenture España, and Forward Deployed AI Engineer at Keyrus with a €70–85K band.
Six job ads with AI Engineer in the title, all open in Spain in the summer of 2026.

These six were open at the same time, all in Spain. One sits inside the client’s office, another automates internal processes, another orchestrates product agents. Similar titles for jobs that are nothing alike, so let’s dig in.

The definition is still forming, so we lean on Andrew Ng’s skills map and on the almost 7,000 job ads Alexey Grigorev has collected, and cross them with what we see hiring in Spain.

This is the first of three pieces on the AI engineer. The other two run on our own data: what an AI Engineer earns in Spain and who holds those jobs today.

What an AI Engineer is

An AI Engineer builds and maintains the systems around a language model. The most common case, and the one behind most job ads, is this example from Yarchi: the legal team has 4,000 contracts and wants to ask questions about them in plain language. Somebody has to chop those contracts up, decide which pieces the model sees for each question, and measure how many answers come back wrong. That is the job.

Andrew Ng would rather talk about AI engineering skills than about the AI Engineer role, because the skills are broader than the job. His comparison is a good one. Today any developer should know how to work with the cloud, and only a few have “Cloud Engineer” in their title. AI is going the same way. Full-stack, data engineers, DevOps, ML engineers and, yes, AI engineers too — all of them are going to need AI engineering skills.

That distinction shows up in our data. In the sample of technical profiles we analysed in Spain, 5,980 declare generative-AI skills, and only 849 of them carry the AI Engineer title. That is 14.2%. There are 1,878 with the title; what the 849 tells you is how many of those also declare the skill. We take it apart in the third piece of this series.

The title covers four dimensions

closest to the model closest to the business model → business

01 Model

ML Engineer

Trains and adapts models: labelling the data, fine-tuning, and judging when that beats calling a large model. The path every roadmap describes and almost nobody hires for.

02 Platform

AI Platform Engineer

The internal tooling for everyone who builds: where the calls go, what they cost, and what happens when they fail.

03 Application

LLM Application Engineer

Product on top of somebody else's model. Never touches it. Builds everything around it so it answers with the company's own data: searching the documents and handing the model only the parts it needs.

04 Business

Forward-Deployed / AI Product Engineer

Inside the client's business. Understanding the process counts for as much as writing the code.

Evaluation

It cuts across all four. It means measuring how often the system answers badly, against a fixed set of cases. Between half and seven in ten Spanish job ads for the role name it, depending on where you draw the line on what counts as evaluation.

Which one you need, using Yarchi’s examples. If the problem is that the model knows nothing about your business, you want 03. If it is that three teams are calling four providers and nobody knows what it costs, 02. If what needs automating is a business process and understanding it counts for as much as coding it, 04. Number 01 is the rarest of the four, and if you genuinely need someone who touches the model’s weights, budget for a long search: only 2.9% of the profiles with this title in Spain declare fine-tuning.

This taxonomy comes from Yarchi’s synthesis of the 4,894-description cut that Alexey Grigorev’s dataset held at the time, collected across the United States, Europe and India. Not all of them carry AI Engineer in the title: the corpus also includes platform and data roles. It is the only public dataset we know of on this question.

Chip Huyen separates AI engineering from ML by two things: adapting models and evaluating them, rather than building them. And she splits adaptation into two paths, depending on whether the weights get touched: prompting on one side, fine-tuning on the other. In Spain the second barely shows up. Of the more than 1,800 profiles carrying the title or a variant, only 2.9% say they can do fine-tuning — retraining the model on your data instead of just instructing it. The role as it exists here lives in the application layer. If post-training genuinely gets cheaper, that 2.9% will climb fast and will need measuring again. When we measured it in August it came out at 0.8%, with a skills dictionary that only looked at English-language tags; counting the Spanish ones too takes it to 2.9%.

For hiring this matters more than it looks. The person who trains models and the person who builds you a production system against your own data are not the same person, and they do not cost the same to find. Ask for the wrong profile and you pay for it in months of searching. The 3% comes from the third piece, counting what skills each profile declares.

What skills it takes, according to Andrew Ng

Above we split the title into four jobs. This is a different cut: the skills you will be asked for in any of the four.

Andrew Ng’s skills map is built on more than 10,000 job ads, dozens of interviews with experts, hiring managers and recruiters, and surveys.

It has four top-level skills:

  1. Building and deploying AI applications.
  2. Software engineering fundamentals. Knowing what the tradeoffs are between cost, scalability, reliability and speed. Andrew Ng contrasts this with vibe coding, where you don’t know what decisions your agent is making for you — usually bad ones, because you don’t know what context to give it.
  3. Coding with agents. Having a good mental model of how they work, knowing their limits, and knowing how much to step in and how much to leave them alone.
  4. Shaping the work. As agents get better at meeting a clear spec, the engineer’s job moves to deciding what goes in that spec. That takes product judgement and an understanding of the business and the customer.

Of the four, this article develops the first. The second and the fourth you would ask of any senior engineer, with AI or without it. The third is very important, and we may develop it another time. The first is the only one that only makes sense when there is a model inside the product, and it is the one that decides who you hire.

In another post, Andrew Ng goes into that first skill in detail and breaks it into six pieces:

AI engineering skills · after Andrew Ng

La rama marcada es la que desarrollamos en este artículo.

Andrew Ng’s premise is that software with AI in it has a less predictable output than normal software. You do not know in advance what an LLM is going to answer. That makes building it far more iterative. You write a piece, inspect it, and decide what to try next, and that sequence depends on what you keep finding. He puts it like this: knowing how to pick the next step well is what lets you build reliable systems out of AI components that are not.

The six pieces of building and deploying AI applications

What an AI Engineer decides inside each one.

LLM fundamentals

Understanding how an LLM tokenises and how it generates output is what tells you when you can trust the model and when it is going to fail.

  • What goes in the context window
  • When to go multimodal
  • Cache hits and the knowledge cut-off date
  • Reasoning effort and sampling parameters
  • When to use tool calling
  • When fine-tuning or self-hosting is actually needed

Grounding the model in data

An LLM without good context is useless. RAG over a vector search was the first attempt, but there are more ways to supply that context now.

  • What goes in the prompt and what the model retrieves on its own
  • Vector index, knowledge graph, or a semantic layer over structured data
  • Turning text, PDF, HTML and images into something the model can read
  • Keeping those pipelines clean and current

Building agentic systems

They range from a workflow with a fixed sequence of calls to a harness where the model decides its own next step.

  • What you chain, what you parallelise, what you solve in code and what with an LLM
  • Which tools it can call: MCP, CLI, sandbox
  • What memory, and how you manage context across long sessions
  • Whether several agents need orchestrating
  • The jump to production: guardrails, adversarial input, exfiltration, governance

Evaluation-driven development

For Andrew Ng this is the trait that most separates someone good at building these systems: running a disciplined loop of evals and error analysis. It is hard to master because the right approach changes with the project, and even with the phase.

  • Reading traces and doing exploratory analysis
  • Deciding what to measure, cross-checked against product and business judgement
  • When to evaluate in code, when with LLM-as-a-judge, and when with a person

Running it in production

Different from classic software because of unpredictability, cost and latency.

  • Observability over real usage, and drift detection
  • Responding to model failures and to prompt injection
  • Regression testing and CI/CD with statistical evaluation, calibrated to what an error costs
  • Cutting cost and latency: model choice, distillation, fine-tuning, simplifying the workflow

Machine learning fundamentals

The mental models for deciding when the system's output is uncertain, and knowing how to use machine learning: a model someone else trained, or one you train yourself.

  • Bias and variance
  • Error analysis
  • Data engineering

If you are interviewing, ask about the fourth one, evaluation-driven development: when your system started failing, how did you measure it, over how many cases, and what changed the number? Someone who has done that work answers with concrete numbers.

What we measure in the next two pieces

The next two pieces run on our own data. In the salary one we found a €19,600 gap for the same title, depending on who posts the ad. And in the profile one, the number we did not expect: 43 people out of 1,878 looking for work.

Data and method

The proprietary data that appears here is measured and explained in the other two pieces: what comes from profiles (the generative-AI pool and the 2.9% fine-tuning figure) in the third, and what comes from job ads (the 476, the €19,600 and how many name evaluation) in the second. The rest is other people’s work and we cite it as such:

  • The definition by model adaptation is Chip Huyen’s, in The AI Engineering Stack and in her book AI Engineering (O’Reilly, 2025). The two paths she describes are prompting and fine-tuning, depending on whether they update the weights.
  • The skills map is Andrew Ng’s and we have not verified it. It comes from more than 10,000 job ads, interviews and surveys. The full map and the part we summarise.
  • The four-jobs taxonomy comes from Yarchi’s synthesis of Alexey Grigorev’s dataset: 4,894 descriptions from builtin.com across Los Angeles, New York, London, Amsterdam, Berlin and India, between February and June 2026. That is the cut the taxonomy was built on; the dataset has kept growing and now stands at 6,964 descriptions through August, which is the version we use in the salary piece. We match neither in sample nor in method: theirs is builtin.com across six markets with skills extracted by a model, ours is LinkedIn Spain counting words. Yarchi warns that keyword counts point in a direction rather than to decimals, and Grigorev says of his own that they are less precise than a strict statistical analysis but representative. Our reading of the dataset.

This article gets updated. When there is new data it gets rewritten at this same URL, rather than published again.

Eusebio GracianinothiringCo-founder

Co-founder de nothiring.

All posts