Skip to work
Alberto Barnabò

Applied AI Engineer — European Central Bank, Frankfurt

AlbertoBarnabò

I take LLM pipelines and agentic applications to production at the European Central Bank — and, on my own time, train small models and publish the evidence.

Portrait of Alberto Barnabò
Frankfurt am Main
11,677
model downloads on Hugging Face
10 / 4
open models / datasets
51★
stars on lazy-cat
35
public repositories on GitHub

Live from the Hugging Face and GitHub APIs, updated daily

01 — Selected work

Five things I built, and the numbers behind them.

Illustration: a laptop searching a product catalogue with no cloud in the loop

01 — Retrieval

2,100 q/s

retrieval on a laptop CPU

E-commerce product search

Self-hosted semantic search that runs on CPU, with no per-query fees.

A retriever and a cross-encoder reranker fine-tuned on Amazon’s ESCI data, so that a query like “pan that doesnt stick eggs” finds the right product. Both run on CPU with no per-query fees; the repository is the whole factory — every training script, eval and chart. Five open artifacts, including a 31 MB static model for edge deployment.

nDCG@10 0.748 held-out · 45 ms rerank · 5 open artifacts · $2.20 total GPU spend

sentence-transformers · cross-encoder · ONNX · model2vec

GitHub ↗Hugging Face ↗Demo ↗2,091 downloads

Fiduciary’s mascot: an owl in a suit at a laptop

02 — Fine-tuning

9,168

downloads on Hugging Face

Fiduciary

A local-first financial advisor: Qwen3-4B, LoRA, and tool use on Apple Silicon.

A LoRA fine-tune of Qwen3-4B-Instruct on MLX that speaks like a senior personal-finance advisor, wrapped in a small agent loop that reads a portfolio file from disk and calls live price and news tools. Nothing leaves the machine. Published as fused 4-bit weights, a 56 MB adapter, and GGUF for Ollama and LM Studio.

Qwen3-4B · LoRA on MLX · 4-bit · ~45 min training on a 16 GB M-series · 3 published formats

MLX · LoRA · Qwen3 · function calling · GGUF

GitHub ↗Hugging Face ↗

Field-F1 · 85 image documents

Qwen2.5‑VL, inaugural run

Document typeDocsField-F1Best model
scontrino6
0.881
modello 7301
0.800
fattura (PDF)9
0.698
bolletta20
0.560
busta paga38
0.517
CU2
0.444
F249
0.343

Exactly one document in 255 runs was extracted perfectly. Italian document AI is not solved.

03 — Document AI

104

typed fields, from 467 observed

BurocrazIA

The first benchmark for Italian document AI.

Invoices, payslips, F24 tax forms, utility bills, receipts, CU and 730 — a domain with no public dataset, benchmark or eval until now. A 104-field typed schema curated from 467 fields observed on real documents; 204 gold-annotated documents (119 derived by construction from FatturaPA XML, 85 annotated and independently re-verified); a hallucination-counting scorer; and an inaugural leaderboard with Qwen2.5-VL. Italian document AI is not solved.

204 gold documents · 3,805 fields re-verified · 1 perfect run in 255 · best payslip F1 0.517

document AI · VLM evaluation · KIE · Qwen2.5‑VL

GitHub ↗Dataset ↗Demo ↗242 downloads

Five synthetic thermal receipts from the US, UK, Germany, Italy and France

04 — Synthetic data

32,000

receipts, five locales

Synthetic receipts for OCR

32,000 receipts whose ground truth is exact by construction.

Public receipt datasets are small, single-locale and labelled by humans who make mistakes. This generator renders thermal receipts across five locales, then produces a photo-degraded twin of each — homography, uneven lighting, thermal fade, JPEG grunge — with every word box mapped through the same transform. Arithmetic is re-checked per receipt and anything that does not add up is rejected.

US · UK · DE · IT · FR · 4.3 GB on the Hub · pixel-exact word boxes · structured KIE fields

OCR · synthetic data · Pillow · PyArrow

GitHub ↗Dataset ↗1,562 downloads

lazy-cat: a sleeping orange cat

05 — Agent tooling

18.6×

fewer tokens across 17 tasks

lazy-cat

Claude Code skills that stop the agent from over-working.

Two skills that fire at the only two moments that matter: before choosing an approach (is there an API, a package, a one-liner?) and before writing each block (did anyone ask for this?). Measured across 17 benchmark tasks under three conditions each, the same outcomes cost 4,762 tokens instead of 88,655.

88,655 → 4,762 tokens · 17 benchmark tasks · 3 conditions each · 2 skills

Claude Code · agent tooling

GitHub ↗51 stars

02 — Open on Hugging Face

Every model and dataset, with live downloads.

huggingface.co/albertobarnabo
NameKindDownloads
fiduciary-qwen3-4bmodel
8,737
synthetic-receipts-ocrdataset
1,562
ecommerce-product-search-embeddingsmodel
1,069
prompted_tabfactdataset
974
ecommerce-product-search-rerankermodel
515
fiduciary-qwen3-4b-GGUFmodel
431
ecommerce-product-search-embeddings-basemodel
276
burocraziadataset
242
ecommerce-product-search-embeddings-staticmodel
231
esci-product-search-pairsdataset
168
fiduciary-qwen3-4b-loramodel
0
scenesmith-qwen3-4bmodel
0
Live · all-time downloads · revalidated daily14,623

03 — Tools & experiments

The long tail.

About

Italian, in Frankfurt, building AI that has to work on Monday.

I studied computer science and engineering at Politecnico di Milano and finished with a double master’s degree at Xi’an Jiaotong University, where I spent two years and wrote my thesis on language models.

Since 2025 I’ve been an applied AI engineer on the European Central Bank’s internal AI team, where I take LLM pipelines and agentic applications from proof of concept to production: document understanding, retrieval and search over large collections, and the unglamorous parts that make a pipeline hold up — evaluation sets before implementation, cost and latency, error handling and retries. I also present and defend those choices to the business teams that use them.

On my own time I train small models on a laptop, publish what comes out, and write down what I learned — including the results that didn’t work. The projects above are that habit in public.

2025 —

Applied AI Engineer

European Central Bank, Frankfurt

Internal AI team. I take LLM pipelines and agentic applications from proof of concept to production — document understanding, retrieval and search over large internal collections — and build the evaluation sets and benchmarks that decide what ships. Day to day: containerised services on cloud infrastructure, cost and latency, error handling and retries, and explaining and defending technical choices to the business teams that use them.

2022 – 2024

AI Researcher

Xi’an Jiaotong University

NLP and large language models for fact verification over tables — the master’s thesis.

2021 – 2024

M.Sc. Computer Science & Engineering

Politecnico di Milano · Xi’an Jiaotong University

Double-degree programme: one year in Milan, two in Xi’an.

2017 – 2021

B.Sc. Computer Science & Engineering

Politecnico di Milano

Research · Master’s thesis · Xi’an Jiaotong University & Politecnico di Milano · 2024

Large language models for fact-checking over tabular data

How well can a language model read a table and decide whether a claim about it is true? The thesis studies LLM behaviour on TabFact and FEVEROUS and how prompting strategy changes accuracy. The prompted TabFact variants are published as a dataset.

Paper ↗Dataset ↗GitHub ↗

  • Alberto outside the European Central Bank tower in Frankfurt
    Frankfurt — outside the ECB
  • Alberto at Hua Shan, Shaanxi, among red prayer ribbons
    Shaanxi — Hua Shan, during the Xi’an years
  • Alberto at his master’s graduation at Xi’an Jiaotong University
    Xi’an — master’s graduation, 2024
© 2026 Alberto Barnabò · Built with Next.js · Source on GitHubStats live from Hugging Face & GitHub · last build 2026-10-05