A Dubai client asked us to fill one computer vision role. We screened 23 candidates for it. Almost all of them could describe a convolutional architecture, most had a respectable portfolio of notebooks, and roughly a third had published something. Four of them had ever watched their own model run on a camera in a real building and seen it fall apart. That gap — not seniority, not framework familiarity, not the university — predicted everything. Here is the seven-step process I now use, including the four questions that surface it in the first twenty minutes.
Step 1: Decide Which of the Four Computer Vision Jobs You Are Hiring For
“Computer vision engineer” covers four jobs that share almost no daily work. Getting this wrong is the most expensive mistake available to you, because you will pay a premium for the wrong one.
- The researcher. Invents or substantially modifies architectures. Genuinely needed in perhaps one brief in ten. Expensive, and bored by deployment work.
- The applied modeller. Fine-tunes existing detection, segmentation or tracking models on your data. Competent, widely available, and what most people picture.
- The deployment engineer. Gets a model running inside a latency and power budget on specific hardware — quantisation, batching, streaming, the awkward realities of a camera that drops frames. Scarce, and usually the actual bottleneck.
- The data operations engineer. Owns labelling, quality control, the evaluation set and the retraining loop. Almost never advertised for, and the reason most deployments stop improving after month three.
Write down which two of those four your first hire must cover. In my experience with UAE product teams, the honest answer is nearly always the third and fourth, while the job advert on the desk describes the first.
Step 2: Write the Specification Around the Deployment Target
The standard advert lists PyTorch, OpenCV, TensorRT, YOLO and “experience with deep learning”. Every applied candidate in the Emirates matches it, so it filters nobody and attracts everybody.
Replace the framework inventory with four facts, which together do more filtering than a page of requirements:
- The camera. Fixed or moving, indoor or outdoor, resolution, and whether you control its placement.
- The frame rate you actually need. Not the camera’s maximum — the rate the decision requires. Most teams over-specify this by an order of magnitude and buy hardware they did not need.
- The hardware the model must run on. A Jetson-class edge device, an on-premise server, or cloud inference. This single line changes which candidates are qualified.
- The latency budget and what breaks when you miss it. “200 ms or the barrier does not lift in time” is a specification. “Real-time” is not.
Applications drop, often by half. The ones that arrive open with a question about your hardware, which is exactly the sorting signal you want.
Step 3: Source Against the UAE Reality
There is no large pool of product-grade computer vision engineers sitting in Dubai waiting for a job advert. There are four smaller pools, and they behave differently:
- Security and surveillance analytics. The deepest local pool by some distance, because the region has invested heavily in camera infrastructure. These engineers have shipped onto real hardware in bad conditions, which is exactly the scar tissue you want. Many undersell themselves because the work is unglamorous.
- Industrial and energy inspection. Defect detection, thermal imaging, drone-captured inspection of assets. Strong on edge constraints and on the discipline of a controlled evaluation set.
- Retail and mobility analytics. Footfall, queue measurement, shelf monitoring, parking and access control. Good on tracking and re-identification, and used to privacy constraints.
- Regional university labs. Strong on modelling, typically no deployment experience at all. Fine as a second hire, risky as a first.
Relocating candidates remain a real option, but budget four to ten weeks for employment visa processing, while a candidate already in the UAE on a transferable visa can often start in two. That timing difference decides more searches than capability does. If you are weighing local versus offshore delivery for the wider team, our guide on hiring embedded and IoT firmware engineers in Dubai covers the adjacent profile you will need once cameras meet devices.
Want the four candidates instead of the twenty-three?
We run the four-question screen below before anyone reaches your calendar, so your first call is with someone who has already watched a model fail on a real camera.
Lance-toi — start with a vetted shortlistStep 4: The 20-Minute Screen — Four Questions
This is the part that saved the search. Twenty minutes, four questions, asked in this order. None of them is a trick, and all four are impossible to answer convincingly without having done the work.
Question 1: “Tell me about your worst false positive in production.”
The word “worst” is doing the work. Someone who has deployed will go straight to a specific, slightly embarrassing story: a reflection in a glass door counted as a person for three weeks, a shadow classified as a spill, a mannequin that inflated the footfall numbers for a whole mall campaign. They will tell you how they found out — usually because a human complained, not because a metric moved.
A candidate who has only worked on datasets answers in the abstract: precision and recall trade-offs, the confusion matrix, threshold tuning. All correct. None of it is an answer to the question asked.
Question 2: “Who labelled your data, and what did you do about disagreements?”
Labelling is where applied computer vision actually lives, and almost nobody prepares for a question about it. Strong answers describe a real process: a written labelling guide, a sample double-labelled to measure agreement, an escalation path for ambiguous frames, and at least one painful moment where the guide had to change and the old labels had to be redone.
Weak answers are some version of “we used a public dataset” or “an intern did it”. That is not disqualifying on its own, but it tells you the person has not owned the loop that determines whether your system improves after launch.
Question 3: “What happens to your model at night?”
My favourite question, because it is disarmingly simple and enormously diagnostic. Anyone who has run a camera system knows that night is a different problem: infrared mode changes the colour space entirely, headlights and floodlights blow out regions of the frame, and a model that was fine at noon can degrade sharply after dark.
Engineers who have shipped answer immediately and with irritation, because it has cost them a weekend. They mention IR cut filters, separate day and night evaluation sets, or training with the specific camera’s night output rather than synthetically darkened images. Engineers who have not shipped say “we would need to retrain on night data”, which is the correct sentence and carries no evidence behind it.
Question 4: “How did you know it was working after launch?”
This separates engineers from demo builders more reliably than any technical question I have used. The honest, experienced answer is usually uncomfortable: a sampled manual review every week, a dashboard somebody actually looked at, a complaint channel from the operations team, a drift check when a camera was repositioned.
The answer to be wary of is a benchmark score with no post-launch measurement behind it. A model with no production feedback loop is not a product, and the person who built it has not yet learned the part of the job that matters most.
Of 23 candidates, 19 gave textbook answers to all four questions. Four gave stories. We hired from the four, and the interview beyond that point was mostly confirmation.
Step 5: A Paid Take-Home on Your Own Footage
Skip public datasets entirely. Take a two-minute clip from your actual cameras — deliberately including the worst lighting you have — and a hundred or so labelled frames. Ask for four hours of work, and pay for them.
Be explicit that you are not grading accuracy. You are grading the error analysis. The deliverable you want is a short note saying: here is where it fails, here is my hypothesis for why, here is what I would collect next, and here is what I would not bother fixing. That note is the job.
Three things reliably show up in the strong submissions. They inspect the data before modelling and tell you something about your own footage you did not know. They notice the class imbalance or the labelling inconsistency you accidentally left in. And they say plainly which part of the problem is not worth solving, which is a seniority signal that no take-home graded on a score can capture. Our wider notes on hiring AI and ML engineers in Dubai cover how to keep exercises like this short enough that good candidates actually complete them.
Step 6: Test the UAE Failure Modes Explicitly
This is the step that does not appear in any generic hiring guide, and it is the one that has caught the most expensive surprises for teams here.
Models trained and benchmarked on public data collected mostly in Europe and North America meet conditions in the Emirates that they were never evaluated against:
- Glare and extreme dynamic range. A single outdoor frame can contain deep shade and direct midday sun. Detection quality varies by time of day in a way that fools a single-threshold evaluation.
- Heat haze. For much of the year, objects at distance shimmer and distort on outdoor cameras. Tracking identities across that is genuinely hard, and it is not in anybody’s benchmark.
- Dust. On the lens and in the air. Accuracy degrades gradually rather than failing outright, which means nobody notices until someone compares two months of numbers.
- Flowing garments. Person detection, pose estimation and re-identification models trained largely on Western clothing degrade on the kandura and abaya, because assumptions about silhouette and limb visibility no longer hold.
- Bilingual scene text. Any OCR component must handle Arabic and English, frequently on the same sign, with right-to-left rendering.
- Seasonal density shifts. Crowd patterns during Ramadan and the summer months differ enough to move the operating point of a counting or queue system.
You do not need the candidate to have solved all six. You need them to react correctly when you raise them. The answer you want is a way of finding out before launch — collect a week of footage across times of day, build a stratified evaluation set, measure per-condition rather than in aggregate. The answer to worry about is confidence that the model will generalise.
There is also a compliance layer that a regionally experienced candidate will raise unprompted. Security and surveillance camera systems in Dubai sit under the remit of the Security Industry Regulatory Agency, so approvals can gate your deployment schedule independently of model readiness. And imagery that identifies individuals is personal data under the UAE federal data protection framework, which brings retention, purpose limitation and residency into your architecture. A candidate who mentions these without prompting has shipped here.
Step 7: Tie the Offer and the First 90 Days to a Measured Number
Write the probation review before you write the offer, and make it a single number measured on your own deployed stream. Not a benchmark. Not a model shipped. Something like: false positives per camera per day, on the four hardest cameras, below a stated threshold, measured across day and night.
This does three useful things at once. It forces you to build the measurement infrastructure in month one, which is the thing teams otherwise defer indefinitely. It gives an honest candidate the information they need to tell you the target is unrealistic — and the good ones will, which is itself valuable. And it makes the review a conversation about a number rather than about a feeling.
On structuring the package itself, base salary understates the cost in the UAE and the allowances move total compensation more than the base does. Our Dubai salary negotiation guide has the current structure. If your roadmap also involves an APAC team, the constraints on the same profile differ enough to plan separately — our colleagues cover the equivalent AI infrastructure hiring process in Singapore.
The Three Mistakes That Cost the Most
Hiring the researcher for the deployment job. You pay a premium, the person is bored by month four, and the latency problem that actually blocks the product is still there. Step 1 exists entirely to prevent this.
Grading a take-home on accuracy. It rewards whoever tuned the longest and tells you nothing about judgement. The error-analysis note in step 5 is the artefact that predicts performance.
Never asking about labelling. Almost every system that stalls three months after launch stalls because nobody owns the data loop. If neither of your first two hires can run it, your accuracy is frozen at launch quality forever.
Ready to run this on a real shortlist?
We source computer vision engineers across the UAE and pre-screen them on the four questions and the error-analysis exercise above — and we will tell you when the role you have written is really a deployment hire.
Lance-toi — brief our Dubai teamFrequently Asked Questions
What does a computer vision engineer cost in Dubai in 2026?
In the roles we have placed and benchmarked, a computer vision engineer with three to six years of applied experience generally sits between 25,000 and 40,000 AED per month in base salary, and a senior engineer who has owned an edge deployment end to end sits between 40,000 and 60,000 AED per month. Candidates with genuine embedded or on-device optimisation experience price above that band because the pool is very small regionally. As with every UAE package, base salary understates the cost: housing allowance typically adds 10,000 to 18,000 AED per month, and you should budget separately for medical cover, annual flights, and end-of-service gratuity. Government, defence and industrial inspection work pays a premium over retail analytics, mostly because of clearance and compliance exposure.
Do I need a PhD-level computer vision researcher?
Almost certainly not, and hiring one for applied work is a common and expensive mistake in the UAE market. Fewer than one in ten of the briefs we see genuinely require novel model architecture. The rest require someone who can fine-tune an existing detection or segmentation model, get it running within a latency budget on specific hardware, build the labelling and evaluation loop, and diagnose why accuracy collapses on your cameras at four in the afternoon. That is an applied engineering job. A researcher hired into it is usually bored within six months and frequently leaves, and meanwhile the deployment problems that actually block the product go unsolved.
Why do models trained on public datasets fail in the UAE?
Because the visual conditions differ from the data the public benchmarks were collected in. Outdoor cameras in the Emirates contend with extreme glare, heat haze that distorts objects at distance for much of the year, fine dust on lenses and in the air, and very high contrast between shaded and sunlit areas of the same frame. Person detection and pose models trained largely on Western clothing degrade on flowing garments such as the kandura and abaya, because the silhouette and limb visibility assumptions no longer hold. Any OCR or scene-text component has to handle Arabic and English, often on the same sign, with right-to-left rendering. None of this appears in a benchmark score, which is why step 5 grades error analysis on your own footage rather than accuracy on a public set.
What regulatory constraints apply to camera systems in Dubai?
Two layers matter for your engineering plan. Security and surveillance camera systems in Dubai fall under the remit of the Security Industry Regulatory Agency, which sets requirements for how such systems are installed and operated, so the deployment schedule of a security-analytics product is often gated by approvals rather than by model readiness. Separately, the UAE federal personal data protection framework governs the processing of personal data, and imagery that identifies individuals is personal data, which brings retention limits, purpose limitation and data residency into scope. The practical hiring consequence is that a candidate who has shipped a camera product in the region will raise these constraints unprompted in interview. One who has only worked on datasets will not, and that difference is a useful signal.

William
Talent Sourcing Expert at HireDeveloper.ae. Runs sourcing and screening for senior and lead engineering roles across the UAE.