AI Data Center Bootcamp · Week 2

Fitting the model

A team research task. Read, argue, then write it down in your own words. Do not use AI to write this: we are asking what your team thinks.

The case

A government agency in Riyadh wants an internal assistant. Around 300 staff will use it to ask questions against the agency's own policy documents, in Arabic and English.

They already own the server: one NVIDIA RTX A6000, 48 GB. There is no budget for another GPU this financial year.

You have been asked whether this can work, and what to run on it.

01

Which model would you serve, and at what precision?

Name a specific model and the exact precision format. Say why this one, and why that format.

0 words

02

Does it fit? Show the arithmetic.

Weights, KV cache, and what else you count. State the maximum context length you would configure and why. Show your working.

0 words

03

At peak, how many staff can be served at the same time?

It is not 300. What is the real limit, what does it depend on, and what would you measure to answer this properly?

0 words

04

What breaks first if usage doubles, and what would staff notice?

"It crashes" is not an answer unless you can say when, and what the logs would show. A deployment that dies and one that gets slow are different, and it matters here.

0 words

05

You are given another $10,000. What would you do with it?

What you change, and what you deliberately leave alone. Justify the spend against a real price, from a real provider, that you looked up.

0 words

Your sources

Everything you used. At least one must be something you found yourselves. If two sources disagreed, say which you believed.

You can submit again if you revise it. We read the most recent one.