- Seattle WA, US Stefano Stefani - Issaquah WA, US Nikhil Kandoi - Seattle WA, US Rama Krishna Sandeep Pokkunuri - Redmond WA, US Kalpesh N. Sutaria - Seattle WA, US Navneet Sabbineni - Seattle WA, US Ganesh Kumar Gella - Redmond WA, US Cheng Ran Li - Bellevue WA, US
International Classification:
G06N 99/00 G06N 5/04
Abstract:
Techniques for hosting adding and warming a host are described. In some instances, a method of determining that at least one group of hosts is to be increased by adding an additional host to the group of hosts; sending a request to the group of hosts for a list of machine learning models loaded per host of the group of hosts; receiving, from each host, the list of loaded machine learning models; loading at least a proper subset of list of loaded machine learning models into random access memory of the at least one group; receiving a request to perform an inference; routing the request to the additional host of the group of hosts; performing an inference using the additional host of the group of hosts; and providing a result of the inference to an external entity is described.