ABSTRACTION LAYERS FOR SCALABLE DISTRIBUTED MACHINE LEARNING

    公开(公告)号:US20220101480A1

    公开(公告)日:2022-03-31

    申请号:US17398295

    申请日:2021-08-10

    Abstract: One embodiment provides for a method of transmitting data between multiple compute nodes of a distributed compute system, the method comprising creating a global view of communication operations to be performed between the multiple compute nodes of the distributed compute system, the global view created using information specific to a machine learning model associated with the distributed compute system; using the global view to determine a communication cost of the communication operations; and automatically determining a number of network endpoints for use in transmitting the data between the multiple compute nodes of the distributed compute system.

Patent Agency Ranking