Immich Machine Learning Hardware: Smart Search and Face Recognition Sizing

Immich ML hardware guide

Immich Machine Learning Hardware: Smart Search and Face Recognition Sizing

Immich machine-learning work is bursty: initial indexing can be demanding, while steady-state use may be much lighter. Size hardware around how quickly you want large imports indexed and whether inference shares the main server.

Quick answer

What to buy for this Immich workload

For a normal family library, CPU or supported integrated acceleration may be enough. Large libraries and aggressive indexing targets can justify a supported GPU, and Immich can also move machine-learning work to a separate machine.

Live Amazon hardware

Current products that fit the categories in this guide

Listings come from this Immich sprint’s dedicated Amazon catalogue. Accessories, barebones mini PCs, ambiguous NAS bundles, wrong-capacity storage, and mismatched GPU models are excluded by the normalizer.

Checking the dedicated Immich catalogue…

Buying decision

Choose the system architecture before chasing specifications

Do not buy a GPU from headline performance alone. Driver support, supported acceleration path, VRAM, idle power, and whether the device can be passed through to Immich matter more than gaming benchmarks.

Interactive sizing

Immich Machine-Learning Hardware Selector

Use your workload to get a hardware tier before comparing products. The result is planning guidance, not a substitute for checking the current Immich compatibility documentation.

Compatibility checklist

Four checks before you purchase

Workload first

Size users, library growth, video share, and background jobs before choosing hardware.

Fast application storage

Keep the database and generated application data on reliable SSD storage where practical.

Protected media capacity

Plan usable capacity after redundancy, reserve space, growth, and backup copies.

Verified acceleration

Confirm the exact hardware, drivers, operating system, and Immich acceleration path before purchase.

01

What Immich machine learning actually does

Immich machine learning powers features such as semantic search and face processing. Those jobs are not continuously maxing out the server, so sizing should focus on how much media must be indexed, how quickly you expect the queue to clear, and what else shares the host.

Machine-learning sizing is fundamentally a queue-time decision. Size for indexing time rather than chasing peak AI benchmarks. A small library can tolerate CPU-only work; larger libraries or tight indexing targets may benefit from a supported accelerator or a separate machine-learning worker. For “What Immich machine learning actually does,” decide whether minutes, hours, or overnight processing is acceptable before assigning money and power to an accelerator.

02

Initial import versus steady-state demand

An initial import can create a large backlog that makes modest hardware look permanently slow. After indexing catches up, daily demand may fall sharply. Measure or estimate the one-time backlog separately from steady-state uploads before buying hardware for the worst hour of ownership.

Initial indexing and steady-state inference are different workloads. Size for indexing time rather than chasing peak AI benchmarks. A small library can tolerate CPU-only work; larger libraries or tight indexing targets may benefit from a supported accelerator or a separate machine-learning worker. “Initial import versus steady-state demand” should be evaluated against both phases so a temporary import backlog does not automatically dictate permanent workstation-class hardware.

03

CPU-only inference

CPU-only inference is the simplest deployment because it avoids GPU drivers and device mapping. It can be entirely reasonable for a small library or a user who is happy to let background jobs run overnight. Simplicity is a legitimate performance feature in a home server.

Supported acceleration is more valuable than theoretical compute. Size for indexing time rather than chasing peak AI benchmarks. A small library can tolerate CPU-only work; larger libraries or tight indexing targets may benefit from a supported accelerator or a separate machine-learning worker. Under “CPU-only inference,” confirm device generation, drivers, container access, and the current Immich acceleration documentation before comparing TOPS or gaming benchmarks.

04

Intel OpenVINO path

Immich documents an OpenVINO acceleration path for supported Intel hardware. That makes some Intel systems attractive when you want machine-learning acceleration without a large discrete GPU, but compatibility still depends on the exact device and software environment.

Remote inference changes the hardware boundary. Size for indexing time rather than chasing peak AI benchmarks. A small library can tolerate CPU-only work; larger libraries or tight indexing targets may benefit from a supported accelerator or a separate machine-learning worker. Use “Intel OpenVINO path” to test whether the main server truly needs an accelerator or whether a separate worker can clear heavy jobs while media and database services remain stable.

05

NVIDIA CUDA path

CUDA can provide strong acceleration on supported NVIDIA hardware. A buyer should check current Immich driver and compute requirements, then choose a card that fits the workload and server rather than defaulting to the newest gaming GPU with the highest benchmark score.

Machine-learning sizing is fundamentally a queue-time decision. Size for indexing time rather than chasing peak AI benchmarks. A small library can tolerate CPU-only work; larger libraries or tight indexing targets may benefit from a supported accelerator or a separate machine-learning worker. For “NVIDIA CUDA path,” decide whether minutes, hours, or overnight processing is acceptable before assigning money and power to an accelerator.

06

AMD and other supported paths

Immich also documents additional acceleration paths beyond Intel and NVIDIA. Support can be device- and platform-specific, so an AMD, ARM, or Rockchip build should be based on the current official matrix and the deployment method you are prepared to maintain.

Initial indexing and steady-state inference are different workloads. Size for indexing time rather than chasing peak AI benchmarks. A small library can tolerate CPU-only work; larger libraries or tight indexing targets may benefit from a supported accelerator or a separate machine-learning worker. “AMD and other supported paths” should be evaluated against both phases so a temporary import backlog does not automatically dictate permanent workstation-class hardware.

07

VRAM, memory, and batch behavior

Memory capacity affects how comfortably the server handles the database, application, and ML jobs together. GPU memory affects supported accelerator workloads differently. Treat system RAM and VRAM as separate resources rather than adding the numbers together.

Supported acceleration is more valuable than theoretical compute. Size for indexing time rather than chasing peak AI benchmarks. A small library can tolerate CPU-only work; larger libraries or tight indexing targets may benefit from a supported accelerator or a separate machine-learning worker. Under “VRAM, memory, and batch behavior,” confirm device generation, drivers, container access, and the current Immich acceleration documentation before comparing TOPS or gaming benchmarks.

08

Remote machine-learning worker

A separate machine-learning worker is useful when the primary Immich server is storage-focused, physically small, or already efficient. It lets a more powerful workstation or mini PC process inference while the main server keeps serving the library and database.

Remote inference changes the hardware boundary. Size for indexing time rather than chasing peak AI benchmarks. A small library can tolerate CPU-only work; larger libraries or tight indexing targets may benefit from a supported accelerator or a separate machine-learning worker. Use “Remote machine-learning worker” to test whether the main server truly needs an accelerator or whether a separate worker can clear heavy jobs while media and database services remain stable.

09

Virtualization and device access

Virtualization adds another compatibility layer. Passing a GPU or iGPU through a hypervisor can work, but the plan should be tested against the host platform before purchase. Direct hardware access is often easier when inference performance and reliability are priorities.

Machine-learning sizing is fundamentally a queue-time decision. Size for indexing time rather than chasing peak AI benchmarks. A small library can tolerate CPU-only work; larger libraries or tight indexing targets may benefit from a supported accelerator or a separate machine-learning worker. For “Virtualization and device access,” decide whether minutes, hours, or overnight processing is acceptable before assigning money and power to an accelerator.

10

Power and thermals for burst workloads

Machine-learning bursts create heat even when average power is modest. A tiny chassis that performs well in a short benchmark may throttle during a long indexing queue. Sustained cooling and fan behavior matter when you are processing tens of thousands of assets.

Initial indexing and steady-state inference are different workloads. Size for indexing time rather than chasing peak AI benchmarks. A small library can tolerate CPU-only work; larger libraries or tight indexing targets may benefit from a supported accelerator or a separate machine-learning worker. “Power and thermals for burst workloads” should be evaluated against both phases so a temporary import backlog does not automatically dictate permanent workstation-class hardware.

11

How fast does indexing need to finish

Ask how long indexing actually needs to take. If overnight is acceptable, a modest accelerator may be enough. If a professional or family workflow expects new media searchable within minutes, faster hardware and more parallel capacity have a clearer practical value.

Supported acceleration is more valuable than theoretical compute. Size for indexing time rather than chasing peak AI benchmarks. A small library can tolerate CPU-only work; larger libraries or tight indexing targets may benefit from a supported accelerator or a separate machine-learning worker. Under “How fast does indexing need to finish,” confirm device generation, drivers, container access, and the current Immich acceleration documentation before comparing TOPS or gaming benchmarks.

12

Upgrade only when the queue proves you need it

Upgrade after identifying a queue, not because an accelerator exists. Immich can run useful machine-learning features without turning the server into an AI workstation. The best purchase is the lowest-complexity hardware that meets your target indexing time and power budget.

Remote inference changes the hardware boundary. Size for indexing time rather than chasing peak AI benchmarks. A small library can tolerate CPU-only work; larger libraries or tight indexing targets may benefit from a supported accelerator or a separate machine-learning worker. Use “Upgrade only when the queue proves you need it” to test whether the main server truly needs an accelerator or whether a separate worker can clear heavy jobs while media and database services remain stable.

Questions people ask

Immich hardware questions for this workload

What hardware does Immich Smart Search use?

Those features use Immich machine-learning services. CPU processing can work, while supported acceleration can shorten indexing time when the library or backlog is large. For machine learning, set a target indexing time and choose the simplest supported CPU, accelerator, or remote-worker arrangement that clears the queue fast enough.

What hardware does Immich facial recognition use?

Those features use Immich machine-learning services. CPU processing can work, while supported acceleration can shorten indexing time when the library or backlog is large. For machine learning, set a target indexing time and choose the simplest supported CPU, accelerator, or remote-worker arrangement that clears the queue fast enough.

Is CPU-only machine learning enough for Immich?

The right answer depends on library size, video share, users, background-job urgency, storage architecture, and the exact hardware configuration. For machine learning, set a target indexing time and choose the simplest supported CPU, accelerator, or remote-worker arrangement that clears the queue fast enough.

Can Immich use Intel OpenVINO?

The right answer depends on library size, video share, users, background-job urgency, storage architecture, and the exact hardware configuration. For machine learning, set a target indexing time and choose the simplest supported CPU, accelerator, or remote-worker arrangement that clears the queue fast enough.

Can Immich use NVIDIA CUDA?

Immich documents a CUDA machine-learning acceleration path for supported NVIDIA hardware. Check the current driver and compute requirements before selecting an exact card. For machine learning, set a target indexing time and choose the simplest supported CPU, accelerator, or remote-worker arrangement that clears the queue fast enough.

How much GPU memory does Immich machine learning need?

VRAM needs depend on the supported model, batch behavior, and concurrent work. Home users should choose enough memory for the documented acceleration path rather than buying the largest gaming card by default. For machine learning, set a target indexing time and choose the simplest supported CPU, accelerator, or remote-worker arrangement that clears the queue fast enough.

Why is initial Immich indexing slower than daily use?

Initial imports create a large queue of jobs that may not represent normal daily demand. Size hardware around the indexing time you are willing to accept rather than assuming that first-day utilization will continue forever. For machine learning, set a target indexing time and choose the simplest supported CPU, accelerator, or remote-worker arrangement that clears the queue fast enough.

Can I put Immich machine learning on a separate server?

Yes. Immich can place machine-learning work on a separate machine, which can be useful when the primary media server is compact, storage-focused, or intentionally low power. For machine learning, set a target indexing time and choose the simplest supported CPU, accelerator, or remote-worker arrangement that clears the queue fast enough.

Does machine learning work inside a VM?

Immich can be deployed with virtualization, but GPU/iGPU passthrough, storage paths, networking, and device permissions add complexity. Validate those layers before buying hardware whose value depends on passthrough. For machine learning, set a target indexing time and choose the simplest supported CPU, accelerator, or remote-worker arrangement that clears the queue fast enough.

When is a GPU upgrade actually worth it for Immich?

The right answer depends on library size, video share, users, background-job urgency, storage architecture, and the exact hardware configuration. For machine learning, set a target indexing time and choose the simplest supported CPU, accelerator, or remote-worker arrangement that clears the queue fast enough.

References and methodology

Verify changing requirements before you buy

Cloudzat separates product classes, rejects accessories and ambiguous configurations, and uses the sprint’s dedicated Amazon catalogue. Prices are displayed only when a current featured offer is returned; otherwise the page links to Amazon buying options without inventing a price.

As an Amazon Associate, Cloudzat may earn from qualifying purchases. Product availability, prices, seller terms, and compatibility can change.

Scroll to Top