| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Ooo I want to do this one.
Moved the comment to the description.
Use af_set_stream like af_set_device. i.e. all functions after calling af_set_stream would use the same stream.
The global state would make life hard (or impossible?) for people wanting to use your library in a multithreaded manner. I think that would be a lot of people!
EDIT: Let this comment also count as a "vote" for this issue!
The global state would make life hard (or impossible?) for people wanting to use your library in a multithreaded manner.
We already have a bunch of global states that need to be sanitized for multi-threaded applications. We are certain this can be done easily.
The second alternative will take a bit of work on our end for the C API. For C++ interface, we can just add a default parameter at the end for every function.
I wanted to put the easily implementable option on the table, but from a technical and aesthetic point of view, I too would vote for the second alternative.
EDIT: Let this comment also count as a "vote" for this issue!
Just to make it more consistent, all voters say "+1 for Option [X]".
My vote:
+1 for Option 2
+1 for Option 2
I'm writing about both sides of the coin here. I'm just giving my opinion and trying to point out the strengths and weaknesses of each.
I like option 1. We could use device manager to control it. Device manager will control creation and deletion. It can also control stream-wide sync rather than device sync. The problem is more lines of code to set it each time.
Option 2 obviously gives the user more control and uses less lines of code. The problem with changing the API (ie simply adding a default parameter and no new functions) is how do we use the stream argument with CPU and with other future backends. We could leave it alone and throw warnings.
And then comes the issue of which stream we launch JIT kernels on. Various arrays may be working of different streams. The only way to control it would be to do a device sync which defeats the purpose of JIT.
I think we could actually do a combination of both. We could set the default stream argument to -1 and have set stream set a current stream to use. That way users wont have to specify stream for every function. If the stream is -1, then use the value from the global current stream.
And then comes the issue of which stream we launch JIT kernels on. Various arrays may be working of different streams. The only way to control it would be to do a device sync which defeats the purpose of JIT.
It is not a problem with JIT. It will be a problem for any function that takes more than one input. We'll have to sync the streams / queues for each of the inputs if the inputs are from different streams. There is no other way.
We could set the default stream argument to -1 and have set stream set a current stream to use. That way users wont have to specify stream for every function. If the stream is -1, then use the value from the global current stream.
Well.. yeah. The proposal never said that the users HAVE to specify the stream. We'll just be adding additional functionality that users can use if they choose to.
+1 for Option 2
Is nt this one still legitimate, at least for afcuda ?
Yes, This is related to streaming data from a source on the host to the device efficiently. This is different from CUDA streams although they will be used here.
In the documentation, I don't see anything like setStream or an extra parameter for a stream in functions. Was this idea abandoned?
@ipostr08 As of now, users can only fetch ArrayFire's CUDA stream. The support to set a user created CUDA stream is not yet available. No this feature is not abandoned. We are adding things (Events in 3.7 release - helps with streams) progressively, in smaller increments. Unfortunately, we don't have a ETA on this feature right now but we will update this issue with updates as we move forward.
@umar456 any news about stream support in AFcuda ?
We have implemented events which should improve support for streams. The main problem is the memory manager and its current implementation and how the streams are associated with each device. We will need to create a context to enable this sort of behavior.
We have implemented events which should improve support for streams. The main problem is the memory manager and its current implementation and how the streams are associated with each device. We will need to create a context to enable this sort of behavior.
Can we get the internal stream that arrayfire uses?
We have implemented events which should improve support for streams. The main problem is the memory manager and its current implementation and how the streams are associated with each device. We will need to create a context to enable this sort of behavior.
Can we get the internal stream that arrayfire uses?
https://arrayfire.org/docs/group__cuda__mat.htm#gaec1dc4c2aa935dc61889f23248c8450d
https://arrayfire.org/docs/interop_cuda.htm please check this interop tutorial on how to use these functions
Any news on this?
My throughput is terrible because my program spends most of the time sending data to and from the GPU. It would be great to be able to do operations while the next data is being send to the GPU.
| Back | FazBrowse Home | New Git URL |
Put the C functions in af/device.h
Some options to consider: