FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

[ENH]: Negative errors in errorbar() · Issue #26187 · matplotlib/matplotlib · GitHub

Repository navigation

[ENH]: Negative errors in errorbar() #26187

Description

Update: See summary in #26187 (comment)


Problem

In #21266 (v3.6) we prohibited negative errors to bring the behavior in line with the documentation.

I've got the request to again support negative errors, so that the error bar can be detached from the marker. There seem to be cases where users exactly want this behavior, e.g.

Proposed solution

I think we should support this somehow. Possibilities:

  • Revert to the old behavior and adjust the documentation accordingly. I'm not clear whether not allowing negative errors was really a conscious decision or whether we were only making the code behave according to the documentation. - Technically, negative errors are perfictly fine. The only justification I see for the positive error limitation is to enforce a more narrow semantics that errors are usually positive and we want to guard users against accidental negative errors.
  • Alternatively, we could introduce an explicit allow_negative_errors flag.

Choosing between those, I'm inclined to say that the flag feels a bit cumbersome and we can assume that users are responsible enough to handle their data without us needing to enforce negative errors. After all, even if they are undesired, it should be relatively easy to spot and find out the reason.

Activity

  1. rcomer commented on Jun 26, 2023

    Member

    I do not use errorbar very often but when I have I found the 2-element errorbars unintuitive to specify. If I want to express something like “y ± standard error” then setting yerr=scalar makes complete sense to me. The other times I’ve used it, I calculated some sort of confidence interval and then need to translate that to (y - ymin, ymax - y) before giving it to Matplotlib. Obviously this isn’t a difficult calculation once you’ve figured out what you need to do, but I would have liked to be able to pass something like yrange=(ymin, ymax).

    Obviously I don’t know anything about the use-case for the requested negative errors, but that way of specifying seems even less intuitive to me. So I wonder if allowing users to directly specify the range would help there too.

    Also appreciate that there are disadvantages in having two conflicting routes to get the same outcome!

  2. oscargus commented on Jun 26, 2023

    Member

    Reading both suggestions, wouldn't a way be to keep behavior as is and add yrange which just ignores everything?

    If it was only a matter of doing yerr = yrange - y it would be one thing, but then currently it is a bit less straightforward than that.

    (I'm not against allowing negative errors either, but I really see the point of yrange as well.)

  3. timhoffm commented on Jun 26, 2023

    MemberAuthor

    Ah, I think I remember the issue with negative errors. When defining the error range (y-neg_error, y+pos_error) via two arguments y, yerr is the minus sign of -neg_error included in the definition or not. IMHO there is no canoncial way to define this, both seem conceptual reasonable. To prevent accidental errors, we enforce both to be positive (and assume the standard case that the lower error limit is smaller than the value.

    So, unconditionally allowing negative errors is out. This leaves us with the flag from the original proposal.


    We cannot cramp the range logic into the yerr parameter. Having mutually exclusive yerr, yrange parameters is usually a code smell. However here, it may actually be a viable option. (Practiality beats purity):

    plt.errorbar(x, y, yerr=...)
    plt.errorbar(x, y, yrange=...)
    

    both feel reasonable.

    The yrange is partly orthogonal to the negative yerr request. However, yrange could be an alternative solution so that one maybe does not need negative errors.

    That said, what is the intended data structure for yrange?

    • (N, 2) array-like
    • tuple of 2 N-arrays (ymin, ymax)
    • both

    I'm inclined to support both.

  4. rcomer commented on Jun 26, 2023

    Member

    That said, what is the intended data structure for yrange?

    • (N, 2) array-like
    • tuple of 2 N-arrays (ymin, ymax)
    • both

    I'm inclined to support both.

    Supporting both could be ambiguous if the user passes two 2-element arrays. I would favour (2, N) array-like, to be consistent with yerr.

  5. oscargus commented on Jun 26, 2023

    Member

    Wouldn't it be possible to support both and explicitly claim the interpretation of a (2, 2) array? (I'm basically saying that since I do not know how to correctly create one or the other... Except for making it an ndarray and transpose if it is wrong...)

  6. timhoffm commented on Jun 26, 2023

    MemberAuthor

    Conceptually I would write the (N=2, 2) array like this in basic data types: [(min0, max0), (min1, max1)], whereas "tuple of 2 2-arrays" would be ([min0, min1], [max0, max1]); i.e. the fixed size is the tuple and the potentially variable quantity is the list. I think it would thus be bearable to say that 2x2 elements where the outer is an tuple is interpreted as tuple of 2 array-likes. - And more general speaking the tuple of ... is more explicit that the "catch-all" array-like and should thus have precedence.

    My expectation is that having separte ymin, ymax arrays is quite common because one may calculate them individually using vectorized numpy operations: e.g. ymin = y - dy_min; ymax = y + dy_max. Of course one can turn them into a single (N, 2) array using np.array([ymin, ymax]).T but that's always a bit cumbersome and hard to read for inexperienced users.

  7. jklymak commented on Jun 26, 2023

    Member

    I don't understand the original use case. I understand the use case for a yrange but it's so trivially the same as yerr that I'm not a huge fan of adding it. I'd personally just let the upper bound be negative if yerr is two elements or has two rows (2, n).

  8. timhoffm commented on Jun 26, 2023

    MemberAuthor

    I'd personally just let the upper bound be negative

    We actually have a conceptual ambiguity. This becomes more obvious in the reverse case that the lower bound is negative:

    errorbar(x, y=[0, 0, 0], yerr=[(-2, 10), (-3, 10), (-4, 10)])
    

    do we/users expect the lower limits of the errorbars to be (-2, -3, -4) or (2, 3, 4). The latter would be the consistent extension to our current interpretation. But the former would be equally intuitive. We therefore decided to just not support negative errors - and by that circumvent the sign ambiguity problem.

  9. jklymak commented on Jun 26, 2023

    Member

    @timhoffm Thanks, I agree that is semi-confusing if the user thinks it is (+/-) and put the minus sign in for the bottom errorbar.

    I guess in that case I lean towards keeping errorbar as-is and asking users who need their error bars to be not centered on their data to make two calls. It's a touch awkward, but I don't think it's as awkward as what we are asking them to do now. eg:

    y = [0, 0, 0]
    yrange = np.array([[2, 10], [3, 10], [4, 10]])
    ymid = np.mean(yrange, axis=0)
    dy = np.diff(yrange) / 2
    
    # markers:
    errorbar(x, y=[0, 0, 0], yerr=[(0, 0), (0, 0), (0, 0)])
    # bars:
    errorbar(x, y=ymin, yerr=[dy, dy])
    

    It's a bit of fussing for the user, but they really are asking to do something strange, and I don't know that we should complicate the library to support it.

  10. oscargus commented on Jun 26, 2023

    Member

    What about legend entries? And having two ErrorbarContainers?

    I guess that the question is if it is common enough to complicate the library or uncommon enough to complicate it for the users...

    Checked the code and it is unfortunately not that straightforward. I was sort of hoping that the negativity check was quite early. In that case one could just do that if x/yerr was set and if not set it to possibly negative values from x/yrange and that is it. But it requires a bit more surgery than that.

  11. jklymak commented on Jun 26, 2023

    Member

    What about legend entries? And having two ErrorbarContainers?

    errorbar(x, y=[0, 0, 0], yerr=[(0, 0), (0, 0), (0, 0)], label='Label Me', color='C0')
    errorbar(x, y=ymin, yerr=[dy, dy], label='_dont_label_me', color='C0')
    

    I don't think it's reasonable to have a turnkey solution for all possible plot permutations. I have never seen an errorbar plot with the data point outside the error bounds. Sure, there are cases where you know the theoretical data spread and want to plot actual data points on top of that spread, but those are really two different data sets, and I don't have a problem asking the user to plot those in two separate calls. And in that case, I'd do

    # markers:
    scatter(x, y=[0, 0, 0], label='Data')
    # bars:
    errorbar(x, y=ymin, yerr=[dy, dy], label='Theory')
    
  12. jklymak commented on Jun 26, 2023

    Member

    Actually it would be really useful to have some context for why a user wants this.

  13. astromancer commented on Jun 27, 2023

    I'm with @jklymak on this one. Errorbars are meant to indicate the the variance around a mean. Having them separate from the markers does not represent something that is physically interpretable without additional information. IMO allowing negative errors will introduce more confusion than benefit.

  14. oscargus commented on Jun 27, 2023

    Member

    Admittedly, I do not understand the negative error use case either, so will be interesting to learn.

    But being able to specify the range is a clear improvement and this would be a way to solve both.

  15. timhoffm commented on Jun 27, 2023

    MemberAuthor

    To summarize:

    • We've found that supporting negative error values is out of question because of semantic ambiguity See [ENH]: Negative errors in errorbar() #26187 (comment).
    • Having yrange as an alternative to yerr is meaningful and viable. - We still have to decide whether this extension is worth it. (As a side effect, this would support ranges outside of the value).

    I'll change the topic to discuss whether we want yrange.


    Original requester feedback: The case that the value is outside the range is a numerical artifact. It should be just on the edge, but numerics pushes it out and as a consequence a previously working plot started to raise with 3.6. - I conclude that negative errors are not really needed (apart from the fact that we would not have a reasonable API to support them).

  16. changed the title [-][ENH]: Negative errors in errorbars[/-] [+][ENH]: errorbar(x, y, yrange=...) as an alternative to errorbar(x, y, yerr=...)[/+] on Jun 27, 2023
  17. jklymak commented on Jun 27, 2023

    Member

    Original requester feedback: The case that the value is outside the range is a numerical artifact. It should be just on the edge, but numerics pushes it out and as a consequence a previously working plot started to raise with 3.6. - I conclude that negative errors are not really needed (apart from the fact that we would not have a reasonable API to support them).

    Well, this is now off topic, but... Do we maybe want to clip and warn rather than error if negative error bars are used? Someone who makes a semantic error will get a bad looking errorbar and a warning. Someone who wants what your original commenter wants will get a warning, but a still usable plot. If they want the warning to go away, they can clip themselves?

  18. changed the title [-][ENH]: errorbar(x, y, yrange=...) as an alternative to errorbar(x, y, yerr=...)[/-] [+][ENH]: Negative errors in errorbar()[/+] on Jun 29, 2023
  19. timhoffm commented on Jun 29, 2023

    MemberAuthor

    Do we maybe want to clip and warn rather than error if negative error bars are used?

    In the face of ambiguity, refuse the temptation to guess.

    The original problem was a numerical edge case. It's not on us to handle this. The sign interpretation issue in yerr is so crititcal that everything but erroring out is prone to misuse.

    Summary: We don't want any change to the current yerr behavior.

    To keep things separate, I've opened a seprate issue for the yrange discussion.

  20. anntzer commented on Jul 1, 2023

    Contributor

    @timhoffm Sorry to be late to the discussion, but I don't really agree there's any ambiguity in errorbar(x, y=[0, 0, 0], yerr=[(-2, 10), (-3, 10), (-4, 10)]) (the example you gave above). If yerr=[2, 3, 4] means "symmetric errors of +/-2, +/-3, +/-4 around each point, then it seems like normal numpy semantics that yerr=[(2, 2), (3, 3), (4, 4)] also means the same thing (*). If you agree with that then yerr=[(-2, x), (-3, y), (-4, z)] means that the lower bounds start at 2, 3, 4 (i.e. the current interpretation).

    I agree that it's not the nicest API, but I don't agree there's ambiguity there (if we want to use the semantically normal API expansion), and would thus just be in favor of supporting negative error values with the semantics above. As you noted in #26187 (comment), there are legitimate cases where, e.g. due to statistical sampling, the errorbar range cannot be absolutely guaranteed to contain the "central" estimate (e.g., bootstrapping-based confidence intervals).

    (*) Well, really it should have been ([2, 3, 4], [2, 3, 4]) to normal match axes broadcasting order, but we're not really consistent in the library wrt. rows vs columns anyways.

    Feel free to re-close if you disagree.

  21. reopened this on Jul 1, 2023
  22. jklymak commented on Jul 1, 2023

    Member

    It's not that the API is ambiguous perse but that an equally reasonable API would have specified the lower bounds as negative offsets. People have tried this and been very surprised when the bars aren't cantered on the main value.

  23. anntzer commented on Jul 1, 2023

    Contributor

    Again, I think it is normal numpy semantics that yerr=[2] and yerr=[(2, 2)] mean the same thing. Obviously we could also choose to deviate from normal numpy semantics, but I don't think both options are equally good.

  24. timhoffm commented on Jul 3, 2023

    MemberAuthor

    @anntzer You are formally right. The behavior is consistent with the numpy broadcasting logic. However

    • When you not come from the broadcasting argument, but directly think about yerr as list of (dy-, dy+), both including and excluding the minus sign into dy are conecptually reasonable. And people are trying both intuitively with unexpected results in one case. We don't necessarily have to, but we can provide better support by erroring on negative dy with a clear error message.
    • Real negative dy are an edge case. Except for the one which triggered this discussion, we did not get any complaints on the behavior.

    In the combination of both, I believe not supporting negative errors in yerr is the better choice.

    Related: I believe a yrange parameter would be a better extension #26220. It's a convenience (but non-negligible because if you have yrange the calculation of yerr with correct signs is not obvious to get right), and it also provides a clean API for the few people with the "mean" value outside the range.

  25. QuLogic commented on Jul 5, 2023

    Member

    Re-closing as not planned to match previous status.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions


      Back | FazBrowse Home | New Git URL