FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

Implement streaming support for all backends · Issue #113 · arrayfire/arrayfire · GitHub

Repository navigation

Implement streaming support for all backends #113

Description

Put the C functions in af/device.h

Some options to consider:

  • OPTION 1: Use af_set_stream like af_set_device. i.e. all functions after calling af_set_stream would use the same stream.
  • OPTION 2: Write new functions that take in the stream parameter. For example af_sum_async(&out, in, dim, stream) (or something such) vs af_sum(&out, in, dim)

Activity

  1. self-assigned this
    on Nov 13, 2014
  2. umar456 commented on Nov 13, 2014

    Member

    Ooo I want to do this one.

  3. assigned and unassigned on Nov 13, 2014
  4. pavanky commented on Nov 13, 2014

    MemberAuthor

    Moved the comment to the description.

  5. lukedodd commented on Nov 13, 2014

    @pavanky

    Use af_set_stream like af_set_device. i.e. all functions after calling af_set_stream would use the same stream.

    The global state would make life hard (or impossible?) for people wanting to use your library in a multithreaded manner. I think that would be a lot of people!

    EDIT: Let this comment also count as a "vote" for this issue!

  6. pavanky commented on Nov 13, 2014

    MemberAuthor

    The global state would make life hard (or impossible?) for people wanting to use your library in a multithreaded manner.

    We already have a bunch of global states that need to be sanitized for multi-threaded applications. We are certain this can be done easily.

    The second alternative will take a bit of work on our end for the C API. For C++ interface, we can just add a default parameter at the end for every function.

    I wanted to put the easily implementable option on the table, but from a technical and aesthetic point of view, I too would vote for the second alternative.

    EDIT: Let this comment also count as a "vote" for this issue!

    Just to make it more consistent, all voters say "+1 for Option [X]".

    My vote:

    +1 for Option 2

  7. 9prady9 commented on Nov 13, 2014

    Member

    +1 for Option 2

  8. shehzan10 commented on Nov 27, 2014

    Member

    I'm writing about both sides of the coin here. I'm just giving my opinion and trying to point out the strengths and weaknesses of each.

    I like option 1. We could use device manager to control it. Device manager will control creation and deletion. It can also control stream-wide sync rather than device sync. The problem is more lines of code to set it each time.

    Option 2 obviously gives the user more control and uses less lines of code. The problem with changing the API (ie simply adding a default parameter and no new functions) is how do we use the stream argument with CPU and with other future backends. We could leave it alone and throw warnings.

    And then comes the issue of which stream we launch JIT kernels on. Various arrays may be working of different streams. The only way to control it would be to do a device sync which defeats the purpose of JIT.

    I think we could actually do a combination of both. We could set the default stream argument to -1 and have set stream set a current stream to use. That way users wont have to specify stream for every function. If the stream is -1, then use the value from the global current stream.

  9. pavanky commented on Nov 27, 2014

    MemberAuthor

    And then comes the issue of which stream we launch JIT kernels on. Various arrays may be working of different streams. The only way to control it would be to do a device sync which defeats the purpose of JIT.

    It is not a problem with JIT. It will be a problem for any function that takes more than one input. We'll have to sync the streams / queues for each of the inputs if the inputs are from different streams. There is no other way.

    We could set the default stream argument to -1 and have set stream set a current stream to use. That way users wont have to specify stream for every function. If the stream is -1, then use the value from the global current stream.

    Well.. yeah. The proposal never said that the users HAVE to specify the stream. We'll just be adding additional functionality that users can use if they choose to.

  10. munnybearz commented on Mar 31, 2015

    Contributor

    +1 for Option 2

  11. changed the title [-]Expose Streams from CUDA and Queues from OpenCL[/-] [+]Implement streaming support for all backends[/+] on Jun 18, 2015
  12. 6 remaining items

  13. removed this from the 3.4.0 milestone on Aug 17, 2016
  14. modified the milestones: v3.5.1, v3.5.0 on May 22, 2017
  15. modified the milestones: v3.6.0, v3.5.1 on Jun 16, 2017
  16. removed this from the v3.6.0 milestone on Feb 27, 2018
  17. WilliamTambellini commented on Nov 20, 2018

    Contributor

    Is nt this one still legitimate, at least for afcuda ?

  18. umar456 commented on Nov 20, 2018

    Member

    Yes, This is related to streaming data from a source on the host to the device efficiently. This is different from CUDA streams although they will be used here.

  19. ipostr08 commented on Nov 12, 2019

    In the documentation, I don't see anything like setStream or an extra parameter for a stream in functions. Was this idea abandoned?

  20. 9prady9 commented on Mar 2, 2020

    Member

    @ipostr08 As of now, users can only fetch ArrayFire's CUDA stream. The support to set a user created CUDA stream is not yet available. No this feature is not abandoned. We are adding things (Events in 3.7 release - helps with streams) progressively, in smaller increments. Unfortunately, we don't have a ETA on this feature right now but we will update this issue with updates as we move forward.

  21. WilliamTambellini commented on Mar 27, 2020

    Contributor

    Seen at GDC 2020:

    @umar456 any news about stream support in AFcuda ?

  22. umar456 commented on Mar 27, 2020

    Member

    We have implemented events which should improve support for streams. The main problem is the memory manager and its current implementation and how the streams are associated with each device. We will need to create a context to enable this sort of behavior.

  23. NOBLES5E commented on Apr 2, 2021

    We have implemented events which should improve support for streams. The main problem is the memory manager and its current implementation and how the streams are associated with each device. We will need to create a context to enable this sort of behavior.

    Can we get the internal stream that arrayfire uses?

  24. 9prady9 commented on Apr 2, 2021

    Member

    We have implemented events which should improve support for streams. The main problem is the memory manager and its current implementation and how the streams are associated with each device. We will need to create a context to enable this sort of behavior.

    Can we get the internal stream that arrayfire uses?

    https://arrayfire.org/docs/group__cuda__mat.htm#gaec1dc4c2aa935dc61889f23248c8450d

    https://arrayfire.org/docs/interop_cuda.htm please check this interop tutorial on how to use these functions

  25. NatanBiesmans commented on Oct 6, 2024

    Any news on this?
    My throughput is terrible because my program spends most of the time sending data to and from the GPU. It would be great to be able to do operations while the next data is being send to the GPU.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions


    Back | FazBrowse Home | New Git URL