Class: Ignis::Collective::Transport::HostStagedTransport

Inherits:
Base
  • Object
show all
Defined in:
lib/nvruby/collective/transport/host_staged_transport.rb

Overview

Host-Staged Transport - Fallback when P2P is not available

Uses host (pinned) memory as an intermediate staging buffer. Data path: GPU_A -> Host -> GPU_B

Slower than P2P but always works regardless of topology. Bandwidth limited by PCIe x2 (upload + download).

Instance Attribute Summary

Attributes inherited from Base

#dst_device, #src_device

Class Method Summary collapse

Instance Method Summary collapse

Methods inherited from Base

#initialize, #ready?, #recv_sync, #send_sync, #synchronize!, #to_s

Constructor Details

This class inherits a constructor from Ignis::Collective::Transport::Base

Class Method Details

.available? ⇒ Boolean

Check if host-staged is available (always true)

Returns:

  • (Boolean) —

    Always true



155
156
157
# File 'lib/nvruby/collective/transport/host_staged_transport.rb', line 155

def self.available?
  true
end

.transport_type ⇒ Symbol

Returns Transport type identifier.

Returns:

  • (Symbol) —

    Transport type identifier



17
18
19
# File 'lib/nvruby/collective/transport/host_staged_transport.rb', line 17

def self.transport_type
  :host_staged
end

Instance Method Details

#copy_async(dst_buffer, src_buffer, size, stream) ⇒ void

This method returns an undefined value.

Copy data via host staging

Parameters:

  • dst_buffer (FFI::Pointer) —

    Destination GPU buffer

  • src_buffer (FFI::Pointer) —

    Source GPU buffer

  • size (Integer) —

    Size in bytes

  • stream (FFI::Pointer) —

    CUDA stream (for async ops)



47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
# File 'lib/nvruby/collective/transport/host_staged_transport.rb', line 47

def copy_async(dst_buffer, src_buffer, size, stream)
  ensure_initialized!

  # Get or allocate staging buffer
  staging = get_staging_buffer(size)

  # Step 1: GPU_src -> Host (async)
  CUDA::RuntimeAPI.cudaSetDevice(@src_device)
  status = CUDA::RuntimeAPI.cudaMemcpyAsync(
    staging,
    src_buffer,
    size,
    CUDA::RuntimeAPI::MEMCPY_DEVICE_TO_HOST,
    stream
  )
  CUDA::RuntimeAPI.check_status!(status, "Host-staged D2H copy")

  # Synchronize to ensure data is in host memory
  sync_stream(stream)

  # Step 2: Host -> GPU_dst (async)
  CUDA::RuntimeAPI.cudaSetDevice(@dst_device)
  status = CUDA::RuntimeAPI.cudaMemcpyAsync(
    dst_buffer,
    staging,
    size,
    CUDA::RuntimeAPI::MEMCPY_HOST_TO_DEVICE,
    stream
  )
  CUDA::RuntimeAPI.check_status!(status, "Host-staged H2D copy")
end

#copy_sync(dst_buffer, src_buffer, size) ⇒ void

This method returns an undefined value.

Synchronous copy

Parameters:

  • dst_buffer (FFI::Pointer) —

    Destination

  • src_buffer (FFI::Pointer) —

    Source

  • size (Integer) —

    Size in bytes



84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
# File 'lib/nvruby/collective/transport/host_staged_transport.rb', line 84

def copy_sync(dst_buffer, src_buffer, size)
  ensure_initialized!

  staging = get_staging_buffer(size)

  # GPU_src -> Host
  CUDA::RuntimeAPI.cudaSetDevice(@src_device)
  status = CUDA::RuntimeAPI.cudaMemcpy(
    staging,
    src_buffer,
    size,
    CUDA::RuntimeAPI::MEMCPY_DEVICE_TO_HOST
  )
  CUDA::RuntimeAPI.check_status!(status, "Host-staged D2H sync")

  # Host -> GPU_dst
  CUDA::RuntimeAPI.cudaSetDevice(@dst_device)
  status = CUDA::RuntimeAPI.cudaMemcpy(
    dst_buffer,
    staging,
    size,
    CUDA::RuntimeAPI::MEMCPY_HOST_TO_DEVICE
  )
  CUDA::RuntimeAPI.check_status!(status, "Host-staged H2D sync")
end

#destroy! ⇒ void

This method returns an undefined value.

Clean up staging buffers



161
162
163
164
165
166
167
# File 'lib/nvruby/collective/transport/host_staged_transport.rb', line 161

def destroy!
  @staging_buffers.each_value do |buf|
    CUDA::RuntimeAPI.cudaFreeHost(buf) rescue nil
  end
  @staging_buffers.clear
  @initialized = false
end

#estimated_bandwidth ⇒ Float

Returns Estimated bandwidth (GB/s).

Returns:

  • (Float) —

    Estimated bandwidth (GB/s)



22
23
24
# File 'lib/nvruby/collective/transport/host_staged_transport.rb', line 22

def estimated_bandwidth
  12.0  # PCIe 4.0 x16 / 2 (round trip)
end

#estimated_latency ⇒ Float

Returns Estimated latency (microseconds).

Returns:

  • (Float) —

    Estimated latency (microseconds)



27
28
29
# File 'lib/nvruby/collective/transport/host_staged_transport.rb', line 27

def estimated_latency
  25.0  # Higher due to double copy
end

#initialize! ⇒ void

This method returns an undefined value.

Initialize the transport



33
34
35
36
37
38
39
# File 'lib/nvruby/collective/transport/host_staged_transport.rb', line 33

def initialize!
  return if @initialized

  CUDA::RuntimeAPI.ensure_loaded!
  @staging_buffers = {}  # size -> pinned host buffer
  @initialized = true
end

#recv_async(buffer, staging, size, stream) ⇒ void

This method returns an undefined value.

Async receive (host staging to GPU)

Parameters:

  • buffer (FFI::Pointer) —

    Destination GPU buffer

  • staging (FFI::Pointer) —

    Host staging buffer with data

  • size (Integer) —

    Size in bytes

  • stream (FFI::Pointer) —

    CUDA stream



139
140
141
142
143
144
145
146
147
148
149
150
151
# File 'lib/nvruby/collective/transport/host_staged_transport.rb', line 139

def recv_async(buffer, staging, size, stream)
  ensure_initialized!

  CUDA::RuntimeAPI.cudaSetDevice(@dst_device)
  status = CUDA::RuntimeAPI.cudaMemcpyAsync(
    buffer,
    staging,
    size,
    CUDA::RuntimeAPI::MEMCPY_HOST_TO_DEVICE,
    stream
  )
  CUDA::RuntimeAPI.check_status!(status, "Host-staged recv")
end

#send_async(buffer, size, stream) ⇒ FFI::Pointer

Async send (GPU to host staging)

Parameters:

  • buffer (FFI::Pointer) —

    Source GPU buffer

  • size (Integer) —

    Size in bytes

  • stream (FFI::Pointer) —

    CUDA stream

Returns:

  • (FFI::Pointer) —

    Staging buffer with data



115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
# File 'lib/nvruby/collective/transport/host_staged_transport.rb', line 115

def send_async(buffer, size, stream)
  ensure_initialized!

  staging = get_staging_buffer(size)

  CUDA::RuntimeAPI.cudaSetDevice(@src_device)
  status = CUDA::RuntimeAPI.cudaMemcpyAsync(
    staging,
    buffer,
    size,
    CUDA::RuntimeAPI::MEMCPY_DEVICE_TO_HOST,
    stream
  )
  CUDA::RuntimeAPI.check_status!(status, "Host-staged send")

  staging
end