| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
I do not use errorbar very often but when I have I found the 2-element errorbars unintuitive to specify. If I want to express something like “y ± standard error” then setting yerr=scalar makes complete sense to me. The other times I’ve used it, I calculated some sort of confidence interval and then need to translate that to (y - ymin, ymax - y) before giving it to Matplotlib. Obviously this isn’t a difficult calculation once you’ve figured out what you need to do, but I would have liked to be able to pass something like yrange=(ymin, ymax).
Obviously I don’t know anything about the use-case for the requested negative errors, but that way of specifying seems even less intuitive to me. So I wonder if allowing users to directly specify the range would help there too.
Also appreciate that there are disadvantages in having two conflicting routes to get the same outcome!
Reading both suggestions, wouldn't a way be to keep behavior as is and add yrange which just ignores everything?
If it was only a matter of doing yerr = yrange - y it would be one thing, but then currently it is a bit less straightforward than that.
(I'm not against allowing negative errors either, but I really see the point of yrange as well.)
Ah, I think I remember the issue with negative errors. When defining the error range (y-neg_error, y+pos_error) via two arguments y, yerr is the minus sign of -neg_error included in the definition or not. IMHO there is no canoncial way to define this, both seem conceptual reasonable. To prevent accidental errors, we enforce both to be positive (and assume the standard case that the lower error limit is smaller than the value.
So, unconditionally allowing negative errors is out. This leaves us with the flag from the original proposal.
We cannot cramp the range logic into the yerr parameter. Having mutually exclusive yerr, yrange parameters is usually a code smell. However here, it may actually be a viable option. (Practiality beats purity):
plt.errorbar(x, y, yerr=...) plt.errorbar(x, y, yrange=...)
both feel reasonable.
The yrange is partly orthogonal to the negative yerr request. However, yrange could be an alternative solution so that one maybe does not need negative errors.
That said, what is the intended data structure for yrange?
I'm inclined to support both.
That said, what is the intended data structure for yrange?
- (N, 2) array-like
- tuple of 2 N-arrays (ymin, ymax)
- both
I'm inclined to support both.
Supporting both could be ambiguous if the user passes two 2-element arrays. I would favour (2, N) array-like, to be consistent with yerr.
Wouldn't it be possible to support both and explicitly claim the interpretation of a (2, 2) array? (I'm basically saying that since I do not know how to correctly create one or the other... Except for making it an ndarray and transpose if it is wrong...)
Conceptually I would write the (N=2, 2) array like this in basic data types: [(min0, max0), (min1, max1)], whereas "tuple of 2 2-arrays" would be ([min0, min1], [max0, max1]); i.e. the fixed size is the tuple and the potentially variable quantity is the list. I think it would thus be bearable to say that 2x2 elements where the outer is an tuple is interpreted as tuple of 2 array-likes. - And more general speaking the tuple of ... is more explicit that the "catch-all" array-like and should thus have precedence.
My expectation is that having separte ymin, ymax arrays is quite common because one may calculate them individually using vectorized numpy operations: e.g. ymin = y - dy_min; ymax = y + dy_max. Of course one can turn them into a single (N, 2) array using np.array([ymin, ymax]).T but that's always a bit cumbersome and hard to read for inexperienced users.
I don't understand the original use case. I understand the use case for a yrange but it's so trivially the same as yerr that I'm not a huge fan of adding it. I'd personally just let the upper bound be negative if yerr is two elements or has two rows (2, n).
I'd personally just let the upper bound be negative
We actually have a conceptual ambiguity. This becomes more obvious in the reverse case that the lower bound is negative:
errorbar(x, y=[0, 0, 0], yerr=[(-2, 10), (-3, 10), (-4, 10)])
do we/users expect the lower limits of the errorbars to be (-2, -3, -4) or (2, 3, 4). The latter would be the consistent extension to our current interpretation. But the former would be equally intuitive. We therefore decided to just not support negative errors - and by that circumvent the sign ambiguity problem.
@timhoffm Thanks, I agree that is semi-confusing if the user thinks it is (+/-) and put the minus sign in for the bottom errorbar.
I guess in that case I lean towards keeping errorbar as-is and asking users who need their error bars to be not centered on their data to make two calls. It's a touch awkward, but I don't think it's as awkward as what we are asking them to do now. eg:
y = [0, 0, 0] yrange = np.array([[2, 10], [3, 10], [4, 10]]) ymid = np.mean(yrange, axis=0) dy = np.diff(yrange) / 2 # markers: errorbar(x, y=[0, 0, 0], yerr=[(0, 0), (0, 0), (0, 0)]) # bars: errorbar(x, y=ymin, yerr=[dy, dy])
It's a bit of fussing for the user, but they really are asking to do something strange, and I don't know that we should complicate the library to support it.
What about legend entries? And having two ErrorbarContainers?
I guess that the question is if it is common enough to complicate the library or uncommon enough to complicate it for the users...
Checked the code and it is unfortunately not that straightforward. I was sort of hoping that the negativity check was quite early. In that case one could just do that if x/yerr was set and if not set it to possibly negative values from x/yrange and that is it. But it requires a bit more surgery than that.
What about legend entries? And having two ErrorbarContainers?
errorbar(x, y=[0, 0, 0], yerr=[(0, 0), (0, 0), (0, 0)], label='Label Me', color='C0') errorbar(x, y=ymin, yerr=[dy, dy], label='_dont_label_me', color='C0')
I don't think it's reasonable to have a turnkey solution for all possible plot permutations. I have never seen an errorbar plot with the data point outside the error bounds. Sure, there are cases where you know the theoretical data spread and want to plot actual data points on top of that spread, but those are really two different data sets, and I don't have a problem asking the user to plot those in two separate calls. And in that case, I'd do
# markers: scatter(x, y=[0, 0, 0], label='Data') # bars: errorbar(x, y=ymin, yerr=[dy, dy], label='Theory')
Actually it would be really useful to have some context for why a user wants this.
I'm with @jklymak on this one. Errorbars are meant to indicate the the variance around a mean. Having them separate from the markers does not represent something that is physically interpretable without additional information. IMO allowing negative errors will introduce more confusion than benefit.
Admittedly, I do not understand the negative error use case either, so will be interesting to learn.
But being able to specify the range is a clear improvement and this would be a way to solve both.
To summarize:
I'll change the topic to discuss whether we want yrange.
Original requester feedback: The case that the value is outside the range is a numerical artifact. It should be just on the edge, but numerics pushes it out and as a consequence a previously working plot started to raise with 3.6. - I conclude that negative errors are not really needed (apart from the fact that we would not have a reasonable API to support them).
Original requester feedback: The case that the value is outside the range is a numerical artifact. It should be just on the edge, but numerics pushes it out and as a consequence a previously working plot started to raise with 3.6. - I conclude that negative errors are not really needed (apart from the fact that we would not have a reasonable API to support them).
Well, this is now off topic, but... Do we maybe want to clip and warn rather than error if negative error bars are used? Someone who makes a semantic error will get a bad looking errorbar and a warning. Someone who wants what your original commenter wants will get a warning, but a still usable plot. If they want the warning to go away, they can clip themselves?
Do we maybe want to clip and warn rather than error if negative error bars are used?
In the face of ambiguity, refuse the temptation to guess.
The original problem was a numerical edge case. It's not on us to handle this. The sign interpretation issue in yerr is so crititcal that everything but erroring out is prone to misuse.
Summary: We don't want any change to the current yerr behavior.
To keep things separate, I've opened a seprate issue for the yrange discussion.
@timhoffm Sorry to be late to the discussion, but I don't really agree there's any ambiguity in errorbar(x, y=[0, 0, 0], yerr=[(-2, 10), (-3, 10), (-4, 10)]) (the example you gave above). If yerr=[2, 3, 4] means "symmetric errors of +/-2, +/-3, +/-4 around each point, then it seems like normal numpy semantics that yerr=[(2, 2), (3, 3), (4, 4)] also means the same thing (*). If you agree with that then yerr=[(-2, x), (-3, y), (-4, z)] means that the lower bounds start at 2, 3, 4 (i.e. the current interpretation).
I agree that it's not the nicest API, but I don't agree there's ambiguity there (if we want to use the semantically normal API expansion), and would thus just be in favor of supporting negative error values with the semantics above. As you noted in #26187 (comment), there are legitimate cases where, e.g. due to statistical sampling, the errorbar range cannot be absolutely guaranteed to contain the "central" estimate (e.g., bootstrapping-based confidence intervals).
(*) Well, really it should have been ([2, 3, 4], [2, 3, 4]) to normal match axes broadcasting order, but we're not really consistent in the library wrt. rows vs columns anyways.
Feel free to re-close if you disagree.
It's not that the API is ambiguous perse but that an equally reasonable API would have specified the lower bounds as negative offsets. People have tried this and been very surprised when the bars aren't cantered on the main value.
Again, I think it is normal numpy semantics that yerr=[2] and yerr=[(2, 2)] mean the same thing. Obviously we could also choose to deviate from normal numpy semantics, but I don't think both options are equally good.
@anntzer You are formally right. The behavior is consistent with the numpy broadcasting logic. However
In the combination of both, I believe not supporting negative errors in yerr is the better choice.
Related: I believe a yrange parameter would be a better extension #26220. It's a convenience (but non-negligible because if you have yrange the calculation of yerr with correct signs is not obvious to get right), and it also provides a clean API for the few people with the "mean" value outside the range.
Re-closing as not planned to match previous status.
| Back | FazBrowse Home | New Git URL |
Update: See summary in #26187 (comment)
Problem
In #21266 (v3.6) we prohibited negative errors to bring the behavior in line with the documentation.
I've got the request to again support negative errors, so that the error bar can be detached from the marker. There seem to be cases where users exactly want this behavior, e.g.
Proposed solution
I think we should support this somehow. Possibilities:
Choosing between those, I'm inclined to say that the flag feels a bit cumbersome and we can assume that users are responsible enough to handle their data without us needing to enforce negative errors. After all, even if they are undesired, it should be relatively easy to spot and find out the reason.