6 ms·
One of my jobs at work is to translate data-science-provided AI models, convert them to optimal GPU forms, and send them downstream to DevOps for deployment in
by nwatson 16d ago
One of my jobs at work is to translate data-science-provided AI models, convert them to optimal GPU forms, and send them downstream to DevOps for deployment in a variety of contexts, all contexts adaptable from basic Docker instructions.
I had a couple of very specialized adapters for the first couple of environments that were very aware of the whole GPU conversion and deployment stack (ONNX, TensorRT, Triton, etc.), but even these were very fragile to version upgrades, etc., and required a lot of adaptation from one "stack" to the other. As soon as the frontier models got to be real good I took a more "agnostic" approach. My new framework's goal was to be "mealleable" and very hands-off w.r.t. the details of the conversion / deployment ... after all, these models are very well trained on the whole AI pipeline, including deployment. So now the basis is more or less (a) where are the model files, which Docker image do we want to use, which "conversion / deployment" method(s) do we want to try, what are the inference use cases addressed; (b) what is the conversion / deployment method used? (c) how do we assemble the smoke-test and basic performance test cases for multi-client, single- and batch-processing modes for all use cases? (d) how do we package the end results so that the DevOps person has all they need to make sure they have the proper files and that they can run Docker and check that the inference works? These questions and answers are all wrapped in some very flexible base classes. Every conversion is somewhat different, so each one involves extensive discussions with Claude Code (e.g., "hey, look at this other prior conversion, its raw files, compare with these raw files, develop a conversion / use-case-smoke-and-performance-test, and packaging). Rather than trying to be very hands-on at these lower level, Claude has almost free reign at the "low level" to suggest the best approach, and I so I "talk to an engineer with vast knowledge but (for now) a bit less judgment". With this method though I probably cut the total time to conversion-for-deployment by 80% to 90% now ... the choices and options are vast, and Claude knows a lot more than I do.
At times there is not a ready-made solution for the particular problem at hand, and so then it gets more interesting with how to shoehorn a solution into one of the available technologies. We prefer one of the various builds of the Triton Inference Server (or one of its hardened versions maintained by others).