29.01.2013 Views

Tutorial CUDA

Tutorial CUDA

Tutorial CUDA

SHOW MORE
SHOW LESS

You also want an ePaper? Increase the reach of your titles

YUMPU automatically turns print PDFs into web optimized ePapers that Google loves.

Hierarchy of Concurrent Threads<br />

Threads are grouped into thread blocks<br />

threadID<br />

© NVIDIA Corporation 2008<br />

Kernel = grid of thread blocks<br />

Thread Block 0<br />

…<br />

float x =<br />

input[threadID];<br />

float y = func(x);<br />

output[threadID] = y;<br />

…<br />

Thread Block 1<br />

…<br />

float x =<br />

input[threadID];<br />

float y = func(x);<br />

output[threadID] = y;<br />

…<br />

…<br />

Thread Block N - 1<br />

0 1 2 3 4 5 6 7 0 1 2 3 4 5 6 7 0 1 2 3 4 5 6 7<br />

…<br />

float x =<br />

input[threadID];<br />

float y = func(x);<br />

output[threadID] = y;<br />

…<br />

By definition, threads in the same block may synchronize with<br />

barriers<br />

scratch[threadID] = begin[threadID];<br />

__syncthreads();<br />

int left = scratch[threadID - 1];<br />

Threads<br />

wait at the barrier<br />

until all threads<br />

in the same block<br />

reach the barrier

Hooray! Your file is uploaded and ready to be published.

Saved successfully!

Ooh no, something went wrong!