On Prem - AI Engineer
Putco Inc
- Location
- Des Moines, IA, US
- Track
- AI Infrastructure
- Salary
- $125K–$155K / yr
- Posted
- September 25, 2026
- Source
- Indeed
Job description
Putco is a 50+ year USA manufacturer of truck and SUV accessories (lighting, emblems, trim, OEM programs). We are building a closed, on-prem “mothership” AI stack on our ERP — not a chatbot toy — so the company can measure true books, generate accurate reports, and run catalog/OEM workflows without shipping proprietary data to the open internet.
We already have a first NVIDIA Spark box running private Llama (chat works). We need a full-time owner to harden that stack, connect it deeply to Macola, and lead the upgrade to a ~240B-class multi-GPU workstation so this becomes the operating brain of Putco — not a side project.
What you will own
· Own and harden the closed on-prem LLM environment (NVIDIA Spark today; path to multi-GPU ~240B-class hardware such as 4× RTX 6000 Ada — Western models preferred: Llama / Mistral family).
· Macola (ERP) intelligence: read-only SQL/schema as RAG, accurate report library replacing flawed legacy reports, Excel→Macola BOM / article loaders with human-approve write-back only (AI never silent-writes production data).
· Mothership connectors for the full putco.com catalog + OEM (Ford, GM, Nissan, Mopar): attribution across wholesale / retail / web / OEM; playbooks and catalog generation in sequenced waves.
· LAN-only / firewall-locked deployment (Open WebUI or equivalent); no leaking Putco data to outside AI APIs unless James locks a capped exception path.
· Friday demos to James / Paul with real Macola questions and measurable milestones — not slideware.
· Hand off cleanly from the current contractor SOW; document everything so Putco is not single-threaded on one person again.
What success looks like in 90 days
· Spark (or successor box) is the trusted Macola brain: ai_readonly + schema RAG answering real ops/engineering questions with cited SQL.
· Core backlog reports shipping as Excel + email digests James actually uses.
· BOM / article loader path live under human approval.
· Written hardware + model plan for the ~240B upgrade, priced and sequenced (B&H / CDW class kit).
· Mothership Wave 0–1 underway: full data plane connected and books measurable — not a single-SKU pilot.
Must-haves
· Hands-on experience running open-weight LLMs on NVIDIA hardware (Ollama, vLLM, TensorRT-LLM, or similar) — not only cloud SaaS prompt engineering.
· Strong SQL and willingness to live inside a mature ERP schema (Macola / similar manufacturing ERP is a plus; you will learn ours fast).
· Python (or equivalent) for loaders, RAG pipelines, report automation, and light internal UIs.
· Security mindset: closed LAN, least privilege, no “just paste it into ChatGPT.”
· Builder who ships weekly demos. Comfortable on-site in the Des Moines area with ops, engineering, and sales users.
· Clear English for CEO-level updates.
Nice-to-haves
· Manufacturing / automotive aftermarket / OEM documentation familiarity.
· Experience correcting bad legacy SQL report logic.
· Headless commerce / catalog feed work (Woo / similar) or PLM adjacency (Aletiq is a separate system — do not blend without approval).
· Multi-GPU workstation build / ConnectX / high-VRAM inference tuning.
How we work
· Closed system first. Cloud APIs only under a hard monthly spend cap James sets.
· Paul Elwell owns digital / site execution; you own the Macola + on-prem intelligence layer and how it feeds the mothership.
· Aletiq PLM stays separate unless James expands scope in writing.
Pay: $125,000.00 - $155,000.00 per year
Benefits
- 401(k)
- 401(k) matching
- Dental insurance
- Employee discount
- Flexible spending account
- Health insurance
- Life insurance
- Paid time off
- Vision insurance
Work Location: In person