* [PATCH] xen: fix non-ANSI function warning in irq.c
From: Randy Dunlap @ 2011-01-09 4:00 UTC (permalink / raw)
To: virtualization, xen-devel; +Cc: Jeremy Fitzhardinge
From: Randy Dunlap <randy.dunlap@oracle.com>
Fix sparse warning for non-ANSI function declaration:
arch/x86/xen/irq.c:129:30: warning: non-ANSI function declaration of function 'xen_init_irq_ops'
Signed-off-by: Randy Dunlap <randy.dunlap@oracle.com>
Cc: Jeremy Fitzhardinge <jeremy.fitzhardinge@citrix.com>
---
arch/x86/xen/irq.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
--- lnx0107.orig/arch/x86/xen/irq.c
+++ lnx0107/arch/x86/xen/irq.c
@@ -126,7 +126,7 @@ static const struct pv_irq_ops xen_irq_o
#endif
};
-void __init xen_init_irq_ops()
+void __init xen_init_irq_ops(void)
{
pv_irq_ops = xen_irq_ops;
x86_init.irqs.intr_init = xen_init_IRQ;
^ permalink raw reply
* ICAC2011 deadline extended to Jan 18 (8th International Conference on Autonomic Computing)
From: Ming Zhao @ 2011-01-08 0:47 UTC (permalink / raw)
To: virtualization
========================================================================
Call for Papers
The 8th International Conference on Autonomic Computing
ICAC 2011
http://icac2011.cs.fiu.edu
June 14-18th, 2011 Karlsruhe, Germany
========================================================================
Update:
-------
* Paper submission deadline is extended to: Jan 18 (11:59pm PST).
Scope:
------
ICAC is the leading conference on autonomic computing applications,
technology and foundations. Autonomic computing refers to methods and
means for reducing the human burden of managing computing systems.
Systems introducing new autonomic features are becoming increasingly
prevalent, motivating research that spans a variety of areas, from
computer systems, architecture, databases and networks to machine
learning, control theory, and bio-inspired computing. ICAC brings
together researchers and practitioners across these disciplines to
address the multiple facets of adaptation and self-management in
computing systems and applications from different perspectives.
Autonomic computing solutions are sought for grids, clouds, enterprise
software, data centers, Internet services, embedded systems, and sensor
networks, where resources and applications must be managed to maximize
performance and minimize cost, while maintaining predictable and
reliable behavior in the face of varying workloads, failures, and
malicious threats. Papers are solicited from all areas of autonomic
computing, along three main thrusts:
* Applications of autonomic computing: Systems contributions and
experiences are sought with prototyped or deployed systems and
applications that focus on advancing system independence and increasing
system ability to adapt to an unpredictable environment. Application
areas include but are not limited to:
- Enterprise applications
- Clouds and grids
- Internet services
- Data center or large-scale system management
- Embedded and mobile systems
- Energy management
- Sensor networks, especially issues related to autonomous, distributed
management
- Internet of things
- Other applications of autonomic computing to real problems in science,
engineering, business and society.
* Autonomic computing components and services: Papers are sought that
describe protocols, system-level support, services, or application
components that enhance aspects of system autonomy, self-management,
self-tuning, self-configuration, self-diagnosis, and self-healing, or
improve adaptive capabilities. Examples include:
- Autonomic management of resources, workloads, faults, power/thermal,
and other challenges.
- Management of quality of service, including security and dependability
- Self-managing components, such as servers, storage, network protocols,
or specific application elements
- Monitoring systems for autonomic computing
- Virtual machine, operating systems, hardware or application support
for autonomic computing
- Novel human interfaces for monitoring and controlling autonomic
systems
- Management topics, such as specification and modeling of service-level
agreements, behavior enforcement and tie-in with IT governance.
- Toolkits, frameworks, principles and architectures, from software
engineering practices and experimental methodologies to agent-based
techniques and virtualization.
* Algorithms, theory and foundations of autonomic computing: Analytic
foundations are solicited for building efficient autonomic systems,
predicting their behavior, quantifying their performance, analyzing
their stability, guaranteeing their specifications, or optimizing their
efficacy. These include:
- Decision and analysis techniques and their use, such as machine
learning, control theory, predictive methods, emergent behavior, self-
organizing networks, rule-based systems and bio-inspired techniques
- Fundamental science and theory of self-managing systems:
understanding, controlling or exploiting system behaviors to enforce
autonomic properties
- Algorithms, analysis and theory for performance guarantees
- Foundations of self-diagnostic systems
Papers will be judged on originality, significance, interest,
correctness, clarity and relevance to the broader community. Papers in
the first two thrusts should report on experiences, measurements, user
studies, or other evaluations, as appropriate. Evaluations of a
prototype or large-scale deployment of autonomic systems and
applications is expected. Papers in the third thrust should provide new
fundamental insights into relevant autonomic computing problems.
Full papers (a maximum of 10 pages) and posters (2 pages) are invited on
a wide variety of topics relating to autonomic computing. Submitted
papers must be original work, and may not be under consideration for
another conference or journal. Complete formatting and submission
instructions can be found on the conference web site. Accepted papers
and posters will appear in proceedings distributed at the conference
and available electronically. Authors of accepted papers and posters
are expected to present their work at the conference.
Important dates:
----------------
* Submission Deadline: Jan 18 (11:59pm PST).
* Notification Deadline: March 15th, 2011
* FFinal Manuscript: April 4th, 2011
* Workshop Proposals: October 15th, 2010
Organization:
-------------
* General Chair:
o Hartmut Schmeck, KIT
* Program Chair:
o Joseph Hellerstein, Google
o Tarek Abdelzaher, UIUC
* Industry Chair:
o Eno Thereska, Microsoft Research
* Workshops Chair:
o Tom Holvoet, KU Leuven
* Posters/Demo/Exhibits Chair:
o Michael Beigl, KIT
* Publicity Chair:
o Ming Zhao, Florida International University
* Program Committee:
o Michael Beigl, KIT, Germany
o Umesh Bellur, IIT, India
o Fabian Bustamante, Northwestern University, USA
o Lucy Cherkasova, HP Labs, USA
o Chita Das, Penn State University, USA
o Yixin Diao, IBM Research, USA
o Indranil Gupta, UIUC, USA
o David Hutchison, Lancaster University, UK
o Ravi Iyer, UIUC, USA
o Vana Kalogeraki, Athens University of Economics and Business, Greece
o Jeff Kephart, IBM, USA
o Emre Kiciman, Microsoft Research, USA
o Charles Lefurgy, IBM Research, USA
o Yunhao Liu, HKUST, HK
o Pedro Marron, Duisburg, Germany
o Milan Milenkovic, Intel, US
o Dejan Milojicic, HP Labs, USA
o Priya Narasimhan, CMU, USA
o Manish Parashar, Rutgers University, USA
o Ana Radovanovic, Google, USA
o Anders Robertsson, Lund, Sweden
o Masoud Sadjadi, Florida International University, USA
o Karsten Schwan, Georgia Institute of Technology, USA
o Onn Shehory, IBM Haifa Research Lab, Israel
o Prashant Shenoy, University of Massachusetts, USA
o Sharad Singhal, HP Labs, USA
o Mani Srivastava, UCLA, USA
o Neeraj Suri, TU Darmstadt, Germany
o Eno Thereska, Microsoft Research, UK
o Thiemo Voigt, SICS, Sweden
o Adam Wolisz, TU Berlin, Germany
o Dongyan Xu, Purdue University, USA
o Xiaoyun Zhu, VMware, USA
For more information:
---------------------
Web: http://icac2011.cs.fiu.edu
Email: icac2011@cs.fiu.edu
--
Ming Zhao, Assistant Professor
School of Computing and Information Sciences
Florida International University
Tel: (305) 348-2034, Fax: (305) 348-3549
Web: http://visa.cs.fiu.edu/ming
^ permalink raw reply
* Re: [PATCH 1/1] staging: hv: Removed unneeded call to netif_stop_queue() in hv_netvsc
From: Greg KH @ 2011-01-07 18:08 UTC (permalink / raw)
To: Hank Janssen
Cc: devel@linuxdriverproject.org, Haiyang Zhang,
linux-kernel@vger.kernel.org,
Abhishek Kane (Mindtree Consulting PVT LTD),
virtualization@lists.osdl.org
In-Reply-To: <8AFC7968D54FB448A30D8F38F259C56233F397CF@TK5EX14MBXC114.redmond.corp.microsoft.com>
On Fri, Jan 07, 2011 at 05:48:11PM +0000, Hank Janssen wrote:
>
>
> > -----Original Message-----
> > From: Greg KH [mailto:gregkh@suse.de]
> > Sent: Friday, January 07, 2011 9:32 AM
> >
> > What kernel is this needed for, .37? .38? older than .37?
> >
> > thanks,
> >
> > greg k-h
>
> This is needed for kernels 2.6.37 and newer.
Ok.
> Btw, I have not seen KVP patches that where submitted in linux next yet.
> I have not seen any other requests for changes to them, is there anything else needed for acceptance?
No, they are in my "to-apply" queue, they missed the deadline for .38,
sorry, due to the holiday break. I'll queue them up when .38-rc1 is
out.
thanks,
greg k-h
^ permalink raw reply
* RE: [PATCH 1/1] staging: hv: Removed unneeded call to netif_stop_queue() in hv_netvsc
From: Hank Janssen @ 2011-01-07 17:48 UTC (permalink / raw)
To: Greg KH
Cc: linux-kernel@vger.kernel.org, devel@linuxdriverproject.org,
virtualization@lists.osdl.org,
Abhishek Kane (Mindtree Consulting PVT LTD), Haiyang Zhang
In-Reply-To: <20110107173139.GA12343@suse.de>
> -----Original Message-----
> From: Greg KH [mailto:gregkh@suse.de]
> Sent: Friday, January 07, 2011 9:32 AM
>
> What kernel is this needed for, .37? .38? older than .37?
>
> thanks,
>
> greg k-h
This is needed for kernels 2.6.37 and newer.
Btw, I have not seen KVP patches that where submitted in linux next yet.
I have not seen any other requests for changes to them, is there anything else needed for acceptance?
Thanks,
Hank.
^ permalink raw reply
* Re: [PATCH 1/1] staging: hv: Removed unneeded call to netif_stop_queue() in hv_netvsc
From: Greg KH @ 2011-01-07 17:31 UTC (permalink / raw)
To: Hank Janssen
Cc: linux-kernel, devel, virtualization, Abhishek Kane, Haiyang Zhang
In-Reply-To: <1294421139-18657-1-git-send-email-hjanssen@microsoft.com>
On Fri, Jan 07, 2011 at 09:25:39AM -0800, Hank Janssen wrote:
> From: Hank Janssen <hjanssen@microsoft.com>
>
> Removed the call to netif_stop_queue() in netvsc_probe() as
> the queue is not initialized at that point and further call
> to it after queue initialization is really not necessary.
>
> This change was prompted after an upstream change went into
> 2.6.37 (netif_tx_stop_queue) that now checks if netif_stop_queue
> is called before register with netdev is done.
>
> This will eliminate the warning message to the log when hv_netvsc
> driver starts up.
>
> Signed-off-by: Abhishek Kane <v-abkane@microsoft.com>
> Signed-off-by: Haiyang Zhang <haiyangz@microsoft.com>
> Signed-off-by: Hank Janssen <hjanssen@microsoft.com>
What kernel is this needed for, .37? .38? older than .37?
thanks,
greg k-h
^ permalink raw reply
* [PATCH 1/1] staging: hv: Removed unneeded call to netif_stop_queue() in hv_netvsc
From: Hank Janssen @ 2011-01-07 17:25 UTC (permalink / raw)
To: hjanssen, gregkh, linux-kernel, devel, virtualization
Cc: Abhishek Kane, Haiyang Zhang
From: Hank Janssen <hjanssen@microsoft.com>
Removed the call to netif_stop_queue() in netvsc_probe() as
the queue is not initialized at that point and further call
to it after queue initialization is really not necessary.
This change was prompted after an upstream change went into
2.6.37 (netif_tx_stop_queue) that now checks if netif_stop_queue
is called before register with netdev is done.
This will eliminate the warning message to the log when hv_netvsc
driver starts up.
Signed-off-by: Abhishek Kane <v-abkane@microsoft.com>
Signed-off-by: Haiyang Zhang <haiyangz@microsoft.com>
Signed-off-by: Hank Janssen <hjanssen@microsoft.com>
---
drivers/staging/hv/netvsc_drv.c | 1 -
1 files changed, 0 insertions(+), 1 deletions(-)
diff --git a/drivers/staging/hv/netvsc_drv.c b/drivers/staging/hv/netvsc_drv.c
index 0147b40..54706a1 100644
--- a/drivers/staging/hv/netvsc_drv.c
+++ b/drivers/staging/hv/netvsc_drv.c
@@ -358,7 +358,6 @@ static int netvsc_probe(struct device *device)
/* Set initial state */
netif_carrier_off(net);
- netif_stop_queue(net);
net_device_ctx = netdev_priv(net);
net_device_ctx->device_ctx = device_ctx;
--
1.6.0.2
^ permalink raw reply related
* Re: [PATCH] virtio: remove virtio-pci root device
From: Michael S. Tsirkin @ 2011-01-07 10:24 UTC (permalink / raw)
To: Milton Miller
Cc: Anthony Liguori, Jamie Lokier, Thomas Weber, linux-kernel,
virtualization
In-Reply-To: <virtio-pci-noroot@mdm.bga.com>
On Fri, Jan 07, 2011 at 02:55:06AM -0600, Milton Miller wrote:
> We sometimes need to map between the virtio device and
> the given pci device. One such use is OS installer that
> gets the boot pci device from BIOS and needs to
> find the relevant block device. Since it can't,
> installation fails.
>
> Instead of creating a top-level devices/virtio-pci
> directory, create each device under the corresponding
> pci device node. Symlinks to all virtio-pci
> devices can be found under the pci driver link in
> bus/pci/drivers/virtio-pci/devices, and all virtio
> devices under drivers/bus/virtio/devices.
>
> Signed-off-by: Milton Miller <miltonm@bga.com>
Thanks! I'll try this next week.
Some comments from looking at code:
- This will break apps that look for devices in the old virtio-pci directory.
There are not likely to be many of these,
but better be careful. Let's create a compatibility
option to put symlinks there?
Could be under a compile time option with a deprecation plan.
- This will create ./virtio0 under the pci device, right?
But if the device is then renamed, the installer
won't be able to find it, right?
And need to do pattern matching to find the name is a bit ugly.
Could we create ./virtio/virtio0 under the pci device?
Then one can open virtio directory and then the only file there.
> ---
>
> This is an alternative to the patch by Michael S. Tsirkin
> titled "virtio-pci: add softlinks between virtio and pci"
> https://patchwork.kernel.org/patch/454581/
>
> It creates simpler code, uses less memory, and should
> be even easier use by the installer as it won't have to
> know a virtio symlink to follow (just follow none).
>
> Compile tested only as I don't have kvm setup.
>
>
> diff --git a/drivers/virtio/virtio_pci.c b/drivers/virtio/virtio_pci.c
> index ef8d9d5..4fb5b2b 100644
> --- a/drivers/virtio/virtio_pci.c
> +++ b/drivers/virtio/virtio_pci.c
> @@ -96,11 +96,6 @@ static struct pci_device_id virtio_pci_id_table[] = {
>
> MODULE_DEVICE_TABLE(pci, virtio_pci_id_table);
>
> -/* A PCI device has it's own struct device and so does a virtio device so
> - * we create a place for the virtio devices to show up in sysfs. I think it
> - * would make more sense for virtio to not insist on having it's own device. */
> -static struct device *virtio_pci_root;
> -
> /* Convert a generic virtio device to our structure */
> static struct virtio_pci_device *to_vp_device(struct virtio_device *vdev)
> {
> @@ -629,7 +624,7 @@ static int __devinit virtio_pci_probe(struct pci_dev *pci_dev,
> if (vp_dev == NULL)
> return -ENOMEM;
>
> - vp_dev->vdev.dev.parent = virtio_pci_root;
> + vp_dev->vdev.dev.parent = &pci_dev->dev;
> vp_dev->vdev.dev.release = virtio_pci_release_dev;
> vp_dev->vdev.config = &virtio_pci_config_ops;
> vp_dev->pci_dev = pci_dev;
> @@ -717,17 +712,7 @@ static struct pci_driver virtio_pci_driver = {
>
> static int __init virtio_pci_init(void)
> {
> - int err;
> -
> - virtio_pci_root = root_device_register("virtio-pci");
> - if (IS_ERR(virtio_pci_root))
> - return PTR_ERR(virtio_pci_root);
> -
> - err = pci_register_driver(&virtio_pci_driver);
> - if (err)
> - root_device_unregister(virtio_pci_root);
> -
> - return err;
> + return pci_register_driver(&virtio_pci_driver);
> }
>
> module_init(virtio_pci_init);
> @@ -735,7 +720,6 @@ module_init(virtio_pci_init);
> static void __exit virtio_pci_exit(void)
> {
> pci_unregister_driver(&virtio_pci_driver);
> - root_device_unregister(virtio_pci_root);
> }
>
> module_exit(virtio_pci_exit);
^ permalink raw reply
* [PATCH] virtio: remove virtio-pci root device
From: Milton Miller @ 2011-01-07 8:55 UTC (permalink / raw)
To: Michael S. Tsirkin, Rusty Russell, Anthony Liguori, Jamie Lokier,
Thomas
Cc: Milton Miller
In-Reply-To: <virtio-pci-whynotsimplify@mdm.bga.com>
We sometimes need to map between the virtio device and
the given pci device. One such use is OS installer that
gets the boot pci device from BIOS and needs to
find the relevant block device. Since it can't,
installation fails.
Instead of creating a top-level devices/virtio-pci
directory, create each device under the corresponding
pci device node. Symlinks to all virtio-pci
devices can be found under the pci driver link in
bus/pci/drivers/virtio-pci/devices, and all virtio
devices under drivers/bus/virtio/devices.
Signed-off-by: Milton Miller <miltonm@bga.com>
---
This is an alternative to the patch by Michael S. Tsirkin
titled "virtio-pci: add softlinks between virtio and pci"
https://patchwork.kernel.org/patch/454581/
It creates simpler code, uses less memory, and should
be even easier use by the installer as it won't have to
know a virtio symlink to follow (just follow none).
Compile tested only as I don't have kvm setup.
diff --git a/drivers/virtio/virtio_pci.c b/drivers/virtio/virtio_pci.c
index ef8d9d5..4fb5b2b 100644
--- a/drivers/virtio/virtio_pci.c
+++ b/drivers/virtio/virtio_pci.c
@@ -96,11 +96,6 @@ static struct pci_device_id virtio_pci_id_table[] = {
MODULE_DEVICE_TABLE(pci, virtio_pci_id_table);
-/* A PCI device has it's own struct device and so does a virtio device so
- * we create a place for the virtio devices to show up in sysfs. I think it
- * would make more sense for virtio to not insist on having it's own device. */
-static struct device *virtio_pci_root;
-
/* Convert a generic virtio device to our structure */
static struct virtio_pci_device *to_vp_device(struct virtio_device *vdev)
{
@@ -629,7 +624,7 @@ static int __devinit virtio_pci_probe(struct pci_dev *pci_dev,
if (vp_dev == NULL)
return -ENOMEM;
- vp_dev->vdev.dev.parent = virtio_pci_root;
+ vp_dev->vdev.dev.parent = &pci_dev->dev;
vp_dev->vdev.dev.release = virtio_pci_release_dev;
vp_dev->vdev.config = &virtio_pci_config_ops;
vp_dev->pci_dev = pci_dev;
@@ -717,17 +712,7 @@ static struct pci_driver virtio_pci_driver = {
static int __init virtio_pci_init(void)
{
- int err;
-
- virtio_pci_root = root_device_register("virtio-pci");
- if (IS_ERR(virtio_pci_root))
- return PTR_ERR(virtio_pci_root);
-
- err = pci_register_driver(&virtio_pci_driver);
- if (err)
- root_device_unregister(virtio_pci_root);
-
- return err;
+ return pci_register_driver(&virtio_pci_driver);
}
module_init(virtio_pci_init);
@@ -735,7 +720,6 @@ module_init(virtio_pci_init);
static void __exit virtio_pci_exit(void)
{
pci_unregister_driver(&virtio_pci_driver);
- root_device_unregister(virtio_pci_root);
}
module_exit(virtio_pci_exit);
^ permalink raw reply related
* Re: virtio-pci: add softlinks between virtio and pci
From: Milton Miller @ 2011-01-07 8:54 UTC (permalink / raw)
To: Michael S. Tsirkin, Rusty Russell, Anthony Liguori, Jamie Lokier,
Thomas
In-Reply-To: <20110105191711.GA27489@redhat.com>
On: Wed, 05 Jan 2011 at about 19:17:11 -0000, Michael S. Tsirkin wrote:
> We sometimes need to map between the virtio device and
> the given pci device. One such use is OS installer that
> gets the boot pci device from BIOS and needs to
> find the relevant block device. Since it can't,
> installation fails.
>
> Supply softlinks between these to make it possible.
>
> Gleb, could you please ack that this patch below
> will be enough to fix the installer issue that
> you see?
>
why not remove this device anchor and put the devices
under the pci device? Proposed patch follows as a
reply.
milton
^ permalink raw reply
* Re: Flow Control and Port Mirroring Revisited
From: Simon Horman @ 2011-01-07 1:23 UTC (permalink / raw)
To: Jesse Gross
Cc: Eric Dumazet, Rusty Russell, virtualization, dev, virtualization,
netdev, kvm, Michael S. Tsirkin
In-Reply-To: <AANLkTinJK-nbkP5_ee2cuS8RA7jTB4-bcWmAf4bjSouP@mail.gmail.com>
On Thu, Jan 06, 2011 at 05:38:01PM -0500, Jesse Gross wrote:
[ snip ]
>
> I know that everyone likes a nice netperf result but I agree with
> Michael that this probably isn't the right question to be asking. I
> don't think that socket buffers are a real solution to the flow
> control problem: they happen to provide that functionality but it's
> more of a side effect than anything. It's just that the amount of
> memory consumed by packets in the queue(s) doesn't really have any
> implicit meaning for flow control (think multiple physical adapters,
> all with the same speed instead of a virtual device and a physical
> device with wildly different speeds). The analog in the physical
> world that you're looking for would be Ethernet flow control.
> Obviously, if the question is limiting CPU or memory consumption then
> that's a different story.
Point taken. I will see if I can control CPU (and thus memory) consumption
using cgroups and/or tc.
> This patch also double counts memory, since the full size of the
> packet will be accounted for by each clone, even though they share the
> actual packet data. Probably not too significant here but it might be
> when flooding/mirroring to many interfaces. This is at least fixable
> (the Xen-style accounting through page tracking deals with it, though
> it has its own problems).
Agreed on all counts.
^ permalink raw reply
* Re: Flow Control and Port Mirroring Revisited
From: Jesse Gross @ 2011-01-06 22:38 UTC (permalink / raw)
To: Simon Horman
Cc: Eric Dumazet, Rusty Russell, virtualization, dev, virtualization,
netdev, kvm, Michael S. Tsirkin
In-Reply-To: <20110106124439.GA17004@verge.net.au>
On Thu, Jan 6, 2011 at 7:44 AM, Simon Horman <horms@verge.net.au> wrote:
> On Thu, Jan 06, 2011 at 11:22:42AM +0100, Eric Dumazet wrote:
>> Le jeudi 06 janvier 2011 à 18:33 +0900, Simon Horman a écrit :
>> > Hi,
>> >
>> > Back in October I reported that I noticed a problem whereby flow control
>> > breaks down when openvswitch is configured to mirror a port[1].
>> >
>> > I have (finally) looked into this further and the problem appears to relate
>> > to cloning of skbs, as Jesse Gross originally suspected.
>> >
>> > More specifically, in do_execute_actions[2] the first n-1 times that an skb
>> > needs to be transmitted it is cloned first and the final time the original
>> > skb is used.
>> >
>> > In the case that there is only one action, which is the normal case, then
>> > the original skb will be used. But in the case of mirroring the cloning
>> > comes into effect. And in my case the cloned skb seems to go to the (slow)
>> > eth1 interface while the original skb goes to the (fast) dummy0 interface
>> > that I set up to be a mirror. The result is that dummy0 "paces" the flow,
>> > and its a cracking pace at that.
>> >
>> > As an experiment I hacked do_execute_actions() to use the original skb
>> > for the first action instead of the last one. In my case the result was
>> > that eth1 "paces" the flow, and things work reasonably nicely.
>> >
>> > Well, sort of. Things work well for non-GSO skbs but extremely poorly for
>> > GSO skbs where only 3 (yes 3, not 3%) end up at the remote host running
>> > netserv. I'm unsure why, but I digress.
>> >
>> > It seems to me that my hack illustrates the point that the flow ends up
>> > being "paced" by one interface. However I think that what would be
>> > desirable is that the flow is "paced" by the slowest link. Unfortunately
>> > I'm unsure how to achieve that.
>> >
>>
>> Hi Simon !
>>
>> "pacing" is done because skb is attached to a socket, and a socket has a
>> limited (but configurable) sndbuf. sk->sk_wmem_alloc is the current sum
>> of all truesize skbs in flight.
>>
>> When you enter something that :
>>
>> 1) Get a clone of the skb, queue the clone to device X
>> 2) queue the original skb to device Y
>>
>> Then : Socket sndbuf is not affected at all by device X queue.
>> This is speed on device Y that matters.
>>
>> You want to get servo control on both X and Y
>>
>> You could try to
>>
>> 1) Get a clone of skb
>> Attach it to socket too (so that socket get a feedback of final
>> orphaning for the clone) with skb_set_owner_w()
>> queue the clone to device X
>>
>> Unfortunatly, stacked skb->destructor() makes this possible only for
>> known destructor (aka sock_wfree())
>
> Hi Eric !
>
> Thanks for the advice. I had thought about the socket buffer but at some
> point it slipped my mind.
>
> In any case the following patch seems to implement the change that I had in
> mind. However my discussions Michael Tsirkin elsewhere in this thread are
> beginning to make me think that think that perhaps this change isn't the
> best solution.
I know that everyone likes a nice netperf result but I agree with
Michael that this probably isn't the right question to be asking. I
don't think that socket buffers are a real solution to the flow
control problem: they happen to provide that functionality but it's
more of a side effect than anything. It's just that the amount of
memory consumed by packets in the queue(s) doesn't really have any
implicit meaning for flow control (think multiple physical adapters,
all with the same speed instead of a virtual device and a physical
device with wildly different speeds). The analog in the physical
world that you're looking for would be Ethernet flow control.
Obviously, if the question is limiting CPU or memory consumption then
that's a different story.
This patch also double counts memory, since the full size of the
packet will be accounted for by each clone, even though they share the
actual packet data. Probably not too significant here but it might be
when flooding/mirroring to many interfaces. This is at least fixable
(the Xen-style accounting through page tracking deals with it, though
it has its own problems).
^ permalink raw reply
* [PATCH 10/36] VIDEO: xen-fb, switch to for_each_console
From: Greg Kroah-Hartman @ 2011-01-06 22:22 UTC (permalink / raw)
To: linux-kernel
Cc: linux-fbdev, xen-devel, Jeremy Fitzhardinge, Greg Kroah-Hartman,
Chris Wright, virtualization, Jiri Slaby
In-Reply-To: <20110106215404.GA30624@kroah.com>
From: Jiri Slaby <jslaby@suse.cz>
Use newly added for_each_console for iterating consoles.
Signed-off-by: Jiri Slaby <jslaby@suse.cz>
Cc: Jeremy Fitzhardinge <jeremy@xensource.com>
Cc: Chris Wright <chrisw@sous-sol.org>
Cc: virtualization@lists.osdl.org
Cc: xen-devel@lists.xensource.com
Cc: linux-fbdev@vger.kernel.org
Signed-off-by: Greg Kroah-Hartman <gregkh@suse.de>
---
drivers/video/xen-fbfront.c | 2 +-
1 files changed, 1 insertions(+), 1 deletions(-)
diff --git a/drivers/video/xen-fbfront.c b/drivers/video/xen-fbfront.c
index 428d273..4abb0b9 100644
--- a/drivers/video/xen-fbfront.c
+++ b/drivers/video/xen-fbfront.c
@@ -492,7 +492,7 @@ xenfb_make_preferred_console(void)
return;
acquire_console_sem();
- for (c = console_drivers; c; c = c->next) {
+ for_each_console(c) {
if (!strcmp(c->name, "tty") && c->index == 0)
break;
}
--
1.7.3.2
^ permalink raw reply related
* Re: Flow Control and Port Mirroring Revisited
From: Simon Horman @ 2011-01-06 22:01 UTC (permalink / raw)
To: Eric Dumazet
Cc: Rusty Russell, virtualization, Jesse Gross, dev, virtualization,
netdev, kvm, Michael S. Tsirkin
In-Reply-To: <1294320498.3074.36.camel@edumazet-laptop>
On Thu, Jan 06, 2011 at 02:28:18PM +0100, Eric Dumazet wrote:
> Le jeudi 06 janvier 2011 à 21:44 +0900, Simon Horman a écrit :
>
> > Hi Eric !
> >
> > Thanks for the advice. I had thought about the socket buffer but at some
> > point it slipped my mind.
> >
> > In any case the following patch seems to implement the change that I had in
> > mind. However my discussions Michael Tsirkin elsewhere in this thread are
> > beginning to make me think that think that perhaps this change isn't the
> > best solution.
> >
> > diff --git a/datapath/actions.c b/datapath/actions.c
> > index 5e16143..505f13f 100644
> > --- a/datapath/actions.c
> > +++ b/datapath/actions.c
> > @@ -384,7 +384,12 @@ static int do_execute_actions(struct datapath *dp, struct sk_buff *skb,
> >
> > for (a = actions, rem = actions_len; rem > 0; a = nla_next(a, &rem)) {
> > if (prev_port != -1) {
> > - do_output(dp, skb_clone(skb, GFP_ATOMIC), prev_port);
> > + struct sk_buff *nskb = skb_clone(skb, GFP_ATOMIC);
> > + if (nskb) {
> > + if (skb->sk)
> > + skb_set_owner_w(nskb, skb->sk);
> > + do_output(dp, nskb, prev_port);
> > + }
> > prev_port = -1;
> > }
> >
> > I got a rather nasty panic without the if (skb->sk),
> > I guess some skbs don't have a socket.
>
> Indeed, some packets are not linked to a socket.
>
> (ARP packets for example)
>
> Sorry, I should have mentioned it :)
Not at all, the occasional panic during hacking is good for the soul.
^ permalink raw reply
* [PATCH] hv: don't enable Scatter/Gather
From: Stephen Hemminger @ 2011-01-06 20:28 UTC (permalink / raw)
To: David Miller; +Cc: netdev, Haiyang Zhang, Mike Sterling, virtualization
In-Reply-To: <1FB5E1D5CA062146B38059374562DF729555F841@TK5EX14MBXC121.redmond.corp.microsoft.com>
The HyperV network driver can do scatter/gather but does not
do checksum offload. Since the network stack requires checksum offload
to do direct transmit (sendfile), netdev_register produces a message
when driver is loaded.
Signed-off-by: Stephen Hemminger <shemminger@vyatta.com>
--- a/drivers/staging/hv/netvsc_drv.c 2011-01-06 12:23:48.143333212 -0800
+++ b/drivers/staging/hv/netvsc_drv.c 2011-01-06 12:24:10.447744121 -0800
@@ -390,8 +390,6 @@ static int netvsc_probe(struct device *d
net->netdev_ops = &device_ops;
/* TODO: Add GSO and Checksum offload */
- net->features = NETIF_F_SG;
-
SET_ETHTOOL_OPS(net, ðtool_ops);
SET_NETDEV_DEV(net, device);
^ permalink raw reply
* Re: Flow Control and Port Mirroring Revisited
From: Eric Dumazet @ 2011-01-06 13:28 UTC (permalink / raw)
To: Simon Horman
Cc: Rusty Russell, virtualization, Jesse Gross, dev, virtualization,
netdev, kvm, Michael S. Tsirkin
In-Reply-To: <20110106124439.GA17004@verge.net.au>
Le jeudi 06 janvier 2011 à 21:44 +0900, Simon Horman a écrit :
> Hi Eric !
>
> Thanks for the advice. I had thought about the socket buffer but at some
> point it slipped my mind.
>
> In any case the following patch seems to implement the change that I had in
> mind. However my discussions Michael Tsirkin elsewhere in this thread are
> beginning to make me think that think that perhaps this change isn't the
> best solution.
>
> diff --git a/datapath/actions.c b/datapath/actions.c
> index 5e16143..505f13f 100644
> --- a/datapath/actions.c
> +++ b/datapath/actions.c
> @@ -384,7 +384,12 @@ static int do_execute_actions(struct datapath *dp, struct sk_buff *skb,
>
> for (a = actions, rem = actions_len; rem > 0; a = nla_next(a, &rem)) {
> if (prev_port != -1) {
> - do_output(dp, skb_clone(skb, GFP_ATOMIC), prev_port);
> + struct sk_buff *nskb = skb_clone(skb, GFP_ATOMIC);
> + if (nskb) {
> + if (skb->sk)
> + skb_set_owner_w(nskb, skb->sk);
> + do_output(dp, nskb, prev_port);
> + }
> prev_port = -1;
> }
>
> I got a rather nasty panic without the if (skb->sk),
> I guess some skbs don't have a socket.
Indeed, some packets are not linked to a socket.
(ARP packets for example)
Sorry, I should have mentioned it :)
^ permalink raw reply
* Re: Flow Control and Port Mirroring Revisited
From: Michael S. Tsirkin @ 2011-01-06 12:47 UTC (permalink / raw)
To: Simon Horman
Cc: Rusty Russell, virtualization, Jesse Gross, dev, virtualization,
netdev, kvm
In-Reply-To: <20110106122859.GA13253@verge.net.au>
On Thu, Jan 06, 2011 at 09:29:02PM +0900, Simon Horman wrote:
> On Thu, Jan 06, 2011 at 02:07:22PM +0200, Michael S. Tsirkin wrote:
> > On Thu, Jan 06, 2011 at 08:30:52PM +0900, Simon Horman wrote:
> > > On Thu, Jan 06, 2011 at 12:27:55PM +0200, Michael S. Tsirkin wrote:
> > > > On Thu, Jan 06, 2011 at 06:33:12PM +0900, Simon Horman wrote:
> > > > > Hi,
> > > > >
> > > > > Back in October I reported that I noticed a problem whereby flow control
> > > > > breaks down when openvswitch is configured to mirror a port[1].
> > > >
> > > > Apropos the UDP flow control. See this
> > > > http://www.spinics.net/lists/netdev/msg150806.html
> > > > for some problems it introduces.
> > > > Unfortunately UDP does not have built-in flow control.
> > > > At some level it's just conceptually broken:
> > > > it's not present in physical networks so why should
> > > > we try and emulate it in a virtual network?
> > > >
> > > >
> > > > Specifically, when you do:
> > > > # netperf -c -4 -t UDP_STREAM -H 172.17.60.218 -l 30 -- -m 1472
> > > > You are asking: what happens if I push data faster than it can be received?
> > > > But why is this an interesting question?
> > > > Ask 'what is the maximum rate at which I can send data with %X packet
> > > > loss' or 'what is the packet loss at rate Y Gb/s'. netperf has
> > > > -b and -w flags for this. It needs to be configured
> > > > with --enable-intervals=yes for them to work.
> > > >
> > > > If you pose the questions this way the problem of pacing
> > > > the execution just goes away.
> > >
> > > I am aware that UDP inherently lacks flow control.
> >
> > Everyone's is aware of that, but this is always followed by a 'however'
> > :).
> >
> > > The aspect of flow control that I am interested in is situations where the
> > > guest can create large amounts of work for the host. However, it seems that
> > > in the case of virtio with vhostnet that the CPU utilisation seems to be
> > > almost entirely attributable to the vhost and qemu-system processes. And
> > > in the case of virtio without vhost net the CPU is used by the qemu-system
> > > process. In both case I assume that I could use a cgroup or something
> > > similar to limit the guests.
> >
> > cgroups, yes. the vhost process inherits the cgroups
> > from the qemu process so you can limit them all.
> >
> > If you are after limiting the max troughput of the guest
> > you can do this with cgroups as well.
>
> Do you mean a CPU cgroup or something else?
net classifier cgroup
> > > Assuming all of that is true then from a resource control problem point of
> > > view, which is mostly what I am concerned about, the problem goes away.
> > > However, I still think that it would be nice to resolve the situation I
> > > described.
> >
> > We need to articulate what's wrong here, otherwise we won't
> > be able to resolve the situation. We are sending UDP packets
> > as fast as we can and some receivers can't cope. Is this the problem?
> > We have made attempts to add a pseudo flow control in the past
> > in an attempt to make UDP on the same host work better.
> > Maybe they help some but they also sure introduce problems.
>
> In the case where port mirroring is not active, which is the
> usual case, to some extent there is flow control in place due to
> (as Eric Dumazet pointed out) the socket buffer.
>
> When port mirroring is activated the flow control operates based
> only on one port - which can't be controlled by the administrator
> in an obvious way.
>
> I think that it would be more intuitive if flow control was
> based on sending a packet to all ports rather than just one.
>
> Though now I think about it some more, perhaps this isn't the best either.
> For instance the case where data was being sent to dummy0 and suddenly
> adding a mirror on eth1 slowed everything down.
>
> So perhaps there needs to be another knob to tune when setting
> up port-mirroring. Or perhaps the current situation isn't so bad.
To understand whether it's bad, you'd need to measure it.
The netperf manual says:
5.2.4 UDP_STREAM
A UDP_STREAM test is similar to a TCP_STREAM test except UDP is used as
the transport rather than TCP.
A UDP_STREAM test has no end-to-end flow control - UDP provides none
and neither does netperf. However, if you wish, you can configure netperf with
--enable-intervals=yes to enable the global command-line -b and -w options to
pace bursts of traffic onto the network.
This has a number of implications.
...
and one of the implications is that the max throughput
might not be reached when you try to send as much data as possible.
It might be confusing that this is what netperf does by default with UDP_STREAM:
if the endpoint is much faster than the network the issue might not appear.
--
MST
^ permalink raw reply
* Re: Flow Control and Port Mirroring Revisited
From: Simon Horman @ 2011-01-06 12:44 UTC (permalink / raw)
To: Eric Dumazet
Cc: Rusty Russell, virtualization, Jesse Gross, dev, virtualization,
netdev, kvm, Michael S. Tsirkin
In-Reply-To: <1294309362.3074.11.camel@edumazet-laptop>
On Thu, Jan 06, 2011 at 11:22:42AM +0100, Eric Dumazet wrote:
> Le jeudi 06 janvier 2011 à 18:33 +0900, Simon Horman a écrit :
> > Hi,
> >
> > Back in October I reported that I noticed a problem whereby flow control
> > breaks down when openvswitch is configured to mirror a port[1].
> >
> > I have (finally) looked into this further and the problem appears to relate
> > to cloning of skbs, as Jesse Gross originally suspected.
> >
> > More specifically, in do_execute_actions[2] the first n-1 times that an skb
> > needs to be transmitted it is cloned first and the final time the original
> > skb is used.
> >
> > In the case that there is only one action, which is the normal case, then
> > the original skb will be used. But in the case of mirroring the cloning
> > comes into effect. And in my case the cloned skb seems to go to the (slow)
> > eth1 interface while the original skb goes to the (fast) dummy0 interface
> > that I set up to be a mirror. The result is that dummy0 "paces" the flow,
> > and its a cracking pace at that.
> >
> > As an experiment I hacked do_execute_actions() to use the original skb
> > for the first action instead of the last one. In my case the result was
> > that eth1 "paces" the flow, and things work reasonably nicely.
> >
> > Well, sort of. Things work well for non-GSO skbs but extremely poorly for
> > GSO skbs where only 3 (yes 3, not 3%) end up at the remote host running
> > netserv. I'm unsure why, but I digress.
> >
> > It seems to me that my hack illustrates the point that the flow ends up
> > being "paced" by one interface. However I think that what would be
> > desirable is that the flow is "paced" by the slowest link. Unfortunately
> > I'm unsure how to achieve that.
> >
>
> Hi Simon !
>
> "pacing" is done because skb is attached to a socket, and a socket has a
> limited (but configurable) sndbuf. sk->sk_wmem_alloc is the current sum
> of all truesize skbs in flight.
>
> When you enter something that :
>
> 1) Get a clone of the skb, queue the clone to device X
> 2) queue the original skb to device Y
>
> Then : Socket sndbuf is not affected at all by device X queue.
> This is speed on device Y that matters.
>
> You want to get servo control on both X and Y
>
> You could try to
>
> 1) Get a clone of skb
> Attach it to socket too (so that socket get a feedback of final
> orphaning for the clone) with skb_set_owner_w()
> queue the clone to device X
>
> Unfortunatly, stacked skb->destructor() makes this possible only for
> known destructor (aka sock_wfree())
Hi Eric !
Thanks for the advice. I had thought about the socket buffer but at some
point it slipped my mind.
In any case the following patch seems to implement the change that I had in
mind. However my discussions Michael Tsirkin elsewhere in this thread are
beginning to make me think that think that perhaps this change isn't the
best solution.
diff --git a/datapath/actions.c b/datapath/actions.c
index 5e16143..505f13f 100644
--- a/datapath/actions.c
+++ b/datapath/actions.c
@@ -384,7 +384,12 @@ static int do_execute_actions(struct datapath *dp, struct sk_buff *skb,
for (a = actions, rem = actions_len; rem > 0; a = nla_next(a, &rem)) {
if (prev_port != -1) {
- do_output(dp, skb_clone(skb, GFP_ATOMIC), prev_port);
+ struct sk_buff *nskb = skb_clone(skb, GFP_ATOMIC);
+ if (nskb) {
+ if (skb->sk)
+ skb_set_owner_w(nskb, skb->sk);
+ do_output(dp, nskb, prev_port);
+ }
prev_port = -1;
}
I got a rather nasty panic without the if (skb->sk),
I guess some skbs don't have a socket.
^ permalink raw reply related
* Re: Flow Control and Port Mirroring Revisited
From: Simon Horman @ 2011-01-06 12:29 UTC (permalink / raw)
To: Michael S. Tsirkin
Cc: Rusty Russell, virtualization, Jesse Gross, dev, virtualization,
netdev, kvm
In-Reply-To: <20110106120722.GD12142@redhat.com>
On Thu, Jan 06, 2011 at 02:07:22PM +0200, Michael S. Tsirkin wrote:
> On Thu, Jan 06, 2011 at 08:30:52PM +0900, Simon Horman wrote:
> > On Thu, Jan 06, 2011 at 12:27:55PM +0200, Michael S. Tsirkin wrote:
> > > On Thu, Jan 06, 2011 at 06:33:12PM +0900, Simon Horman wrote:
> > > > Hi,
> > > >
> > > > Back in October I reported that I noticed a problem whereby flow control
> > > > breaks down when openvswitch is configured to mirror a port[1].
> > >
> > > Apropos the UDP flow control. See this
> > > http://www.spinics.net/lists/netdev/msg150806.html
> > > for some problems it introduces.
> > > Unfortunately UDP does not have built-in flow control.
> > > At some level it's just conceptually broken:
> > > it's not present in physical networks so why should
> > > we try and emulate it in a virtual network?
> > >
> > >
> > > Specifically, when you do:
> > > # netperf -c -4 -t UDP_STREAM -H 172.17.60.218 -l 30 -- -m 1472
> > > You are asking: what happens if I push data faster than it can be received?
> > > But why is this an interesting question?
> > > Ask 'what is the maximum rate at which I can send data with %X packet
> > > loss' or 'what is the packet loss at rate Y Gb/s'. netperf has
> > > -b and -w flags for this. It needs to be configured
> > > with --enable-intervals=yes for them to work.
> > >
> > > If you pose the questions this way the problem of pacing
> > > the execution just goes away.
> >
> > I am aware that UDP inherently lacks flow control.
>
> Everyone's is aware of that, but this is always followed by a 'however'
> :).
>
> > The aspect of flow control that I am interested in is situations where the
> > guest can create large amounts of work for the host. However, it seems that
> > in the case of virtio with vhostnet that the CPU utilisation seems to be
> > almost entirely attributable to the vhost and qemu-system processes. And
> > in the case of virtio without vhost net the CPU is used by the qemu-system
> > process. In both case I assume that I could use a cgroup or something
> > similar to limit the guests.
>
> cgroups, yes. the vhost process inherits the cgroups
> from the qemu process so you can limit them all.
>
> If you are after limiting the max troughput of the guest
> you can do this with cgroups as well.
Do you mean a CPU cgroup or something else?
> > Assuming all of that is true then from a resource control problem point of
> > view, which is mostly what I am concerned about, the problem goes away.
> > However, I still think that it would be nice to resolve the situation I
> > described.
>
> We need to articulate what's wrong here, otherwise we won't
> be able to resolve the situation. We are sending UDP packets
> as fast as we can and some receivers can't cope. Is this the problem?
> We have made attempts to add a pseudo flow control in the past
> in an attempt to make UDP on the same host work better.
> Maybe they help some but they also sure introduce problems.
In the case where port mirroring is not active, which is the
usual case, to some extent there is flow control in place due to
(as Eric Dumazet pointed out) the socket buffer.
When port mirroring is activated the flow control operates based
only on one port - which can't be controlled by the administrator
in an obvious way.
I think that it would be more intuitive if flow control was
based on sending a packet to all ports rather than just one.
Though now I think about it some more, perhaps this isn't the best either.
For instance the case where data was being sent to dummy0 and suddenly
adding a mirror on eth1 slowed everything down.
So perhaps there needs to be another knob to tune when setting
up port-mirroring. Or perhaps the current situation isn't so bad.
^ permalink raw reply
* Re: Flow Control and Port Mirroring Revisited
From: Michael S. Tsirkin @ 2011-01-06 12:07 UTC (permalink / raw)
To: Simon Horman
Cc: Rusty Russell, virtualization, Jesse Gross, dev, virtualization,
netdev, kvm
In-Reply-To: <20110106113052.GA2541@verge.net.au>
On Thu, Jan 06, 2011 at 08:30:52PM +0900, Simon Horman wrote:
> On Thu, Jan 06, 2011 at 12:27:55PM +0200, Michael S. Tsirkin wrote:
> > On Thu, Jan 06, 2011 at 06:33:12PM +0900, Simon Horman wrote:
> > > Hi,
> > >
> > > Back in October I reported that I noticed a problem whereby flow control
> > > breaks down when openvswitch is configured to mirror a port[1].
> >
> > Apropos the UDP flow control. See this
> > http://www.spinics.net/lists/netdev/msg150806.html
> > for some problems it introduces.
> > Unfortunately UDP does not have built-in flow control.
> > At some level it's just conceptually broken:
> > it's not present in physical networks so why should
> > we try and emulate it in a virtual network?
> >
> >
> > Specifically, when you do:
> > # netperf -c -4 -t UDP_STREAM -H 172.17.60.218 -l 30 -- -m 1472
> > You are asking: what happens if I push data faster than it can be received?
> > But why is this an interesting question?
> > Ask 'what is the maximum rate at which I can send data with %X packet
> > loss' or 'what is the packet loss at rate Y Gb/s'. netperf has
> > -b and -w flags for this. It needs to be configured
> > with --enable-intervals=yes for them to work.
> >
> > If you pose the questions this way the problem of pacing
> > the execution just goes away.
>
> I am aware that UDP inherently lacks flow control.
Everyone's is aware of that, but this is always followed by a 'however'
:).
> The aspect of flow control that I am interested in is situations where the
> guest can create large amounts of work for the host. However, it seems that
> in the case of virtio with vhostnet that the CPU utilisation seems to be
> almost entirely attributable to the vhost and qemu-system processes. And
> in the case of virtio without vhost net the CPU is used by the qemu-system
> process. In both case I assume that I could use a cgroup or something
> similar to limit the guests.
cgroups, yes. the vhost process inherits the cgroups
from the qemu process so you can limit them all.
If you are after limiting the max troughput of the guest
you can do this with cgroups as well.
> Assuming all of that is true then from a resource control problem point of
> view, which is mostly what I am concerned about, the problem goes away.
> However, I still think that it would be nice to resolve the situation I
> described.
We need to articulate what's wrong here, otherwise we won't
be able to resolve the situation. We are sending UDP packets
as fast as we can and some receivers can't cope. Is this the problem?
We have made attempts to add a pseudo flow control in the past
in an attempt to make UDP on the same host work better.
Maybe they help some but they also sure introduce problems.
--
MST
^ permalink raw reply
* Re: Flow Control and Port Mirroring Revisited
From: Simon Horman @ 2011-01-06 11:30 UTC (permalink / raw)
To: Michael S. Tsirkin
Cc: Rusty Russell, virtualization, Jesse Gross, dev, virtualization,
netdev, kvm
In-Reply-To: <20110106102755.GC12142@redhat.com>
On Thu, Jan 06, 2011 at 12:27:55PM +0200, Michael S. Tsirkin wrote:
> On Thu, Jan 06, 2011 at 06:33:12PM +0900, Simon Horman wrote:
> > Hi,
> >
> > Back in October I reported that I noticed a problem whereby flow control
> > breaks down when openvswitch is configured to mirror a port[1].
>
> Apropos the UDP flow control. See this
> http://www.spinics.net/lists/netdev/msg150806.html
> for some problems it introduces.
> Unfortunately UDP does not have built-in flow control.
> At some level it's just conceptually broken:
> it's not present in physical networks so why should
> we try and emulate it in a virtual network?
>
>
> Specifically, when you do:
> # netperf -c -4 -t UDP_STREAM -H 172.17.60.218 -l 30 -- -m 1472
> You are asking: what happens if I push data faster than it can be received?
> But why is this an interesting question?
> Ask 'what is the maximum rate at which I can send data with %X packet
> loss' or 'what is the packet loss at rate Y Gb/s'. netperf has
> -b and -w flags for this. It needs to be configured
> with --enable-intervals=yes for them to work.
>
> If you pose the questions this way the problem of pacing
> the execution just goes away.
I am aware that UDP inherently lacks flow control.
The aspect of flow control that I am interested in is situations where the
guest can create large amounts of work for the host. However, it seems that
in the case of virtio with vhostnet that the CPU utilisation seems to be
almost entirely attributable to the vhost and qemu-system processes. And
in the case of virtio without vhost net the CPU is used by the qemu-system
process. In both case I assume that I could use a cgroup or something
similar to limit the guests.
Assuming all of that is true then from a resource control problem point of
view, which is mostly what I am concerned about, the problem goes away.
However, I still think that it would be nice to resolve the situation I
described.
^ permalink raw reply
* Re: Flow Control and Port Mirroring Revisited
From: Michael S. Tsirkin @ 2011-01-06 10:27 UTC (permalink / raw)
To: Simon Horman
Cc: Rusty Russell, virtualization, Jesse Gross, dev, virtualization,
netdev, kvm
In-Reply-To: <20110106093312.GA1564@verge.net.au>
On Thu, Jan 06, 2011 at 06:33:12PM +0900, Simon Horman wrote:
> Hi,
>
> Back in October I reported that I noticed a problem whereby flow control
> breaks down when openvswitch is configured to mirror a port[1].
Apropos the UDP flow control. See this
http://www.spinics.net/lists/netdev/msg150806.html
for some problems it introduces.
Unfortunately UDP does not have built-in flow control.
At some level it's just conceptually broken:
it's not present in physical networks so why should
we try and emulate it in a virtual network?
Specifically, when you do:
# netperf -c -4 -t UDP_STREAM -H 172.17.60.218 -l 30 -- -m 1472
You are asking: what happens if I push data faster than it can be received?
But why is this an interesting question?
Ask 'what is the maximum rate at which I can send data with %X packet
loss' or 'what is the packet loss at rate Y Gb/s'. netperf has
-b and -w flags for this. It needs to be configured
with --enable-intervals=yes for them to work.
If you pose the questions this way the problem of pacing
the execution just goes away.
>
> I have (finally) looked into this further and the problem appears to relate
> to cloning of skbs, as Jesse Gross originally suspected.
>
> More specifically, in do_execute_actions[2] the first n-1 times that an skb
> needs to be transmitted it is cloned first and the final time the original
> skb is used.
>
> In the case that there is only one action, which is the normal case, then
> the original skb will be used. But in the case of mirroring the cloning
> comes into effect. And in my case the cloned skb seems to go to the (slow)
> eth1 interface while the original skb goes to the (fast) dummy0 interface
> that I set up to be a mirror. The result is that dummy0 "paces" the flow,
> and its a cracking pace at that.
>
> As an experiment I hacked do_execute_actions() to use the original skb
> for the first action instead of the last one. In my case the result was
> that eth1 "paces" the flow, and things work reasonably nicely.
>
> Well, sort of. Things work well for non-GSO skbs but extremely poorly for
> GSO skbs where only 3 (yes 3, not 3%) end up at the remote host running
> netserv. I'm unsure why, but I digress.
>
> It seems to me that my hack illustrates the point that the flow ends up
> being "paced" by one interface. However I think that what would be
> desirable is that the flow is "paced" by the slowest link. Unfortunately
> I'm unsure how to achieve that.
What if you have multiple UDP sockets with different targets
in the guest?
> One idea that I had was to skb_get() the original skb each time it is
> cloned - that is easy enough. But unfortunately it seems to me that
> approach would require some sort of callback mechanism in kfree_skb() so
> that the cloned skbs can kfree_skb() the original skb.
>
> Ideas would be greatly appreciated.
>
> [1] http://openvswitch.org/pipermail/dev_openvswitch.org/2010-October/003806.html
> [2] http://openvswitch.org/cgi-bin/gitweb.cgi?p=openvswitch;a=blob;f=datapath/actions.c;h=5e16143ca402f7da0ee8fc18ee5eb16c3b7598e6;hb=HEAD
^ permalink raw reply
* Re: Flow Control and Port Mirroring Revisited
From: Eric Dumazet @ 2011-01-06 10:22 UTC (permalink / raw)
To: Simon Horman
Cc: Rusty Russell, virtualization, Jesse Gross, dev, virtualization,
netdev, kvm, Michael S. Tsirkin
In-Reply-To: <20110106093312.GA1564@verge.net.au>
Le jeudi 06 janvier 2011 à 18:33 +0900, Simon Horman a écrit :
> Hi,
>
> Back in October I reported that I noticed a problem whereby flow control
> breaks down when openvswitch is configured to mirror a port[1].
>
> I have (finally) looked into this further and the problem appears to relate
> to cloning of skbs, as Jesse Gross originally suspected.
>
> More specifically, in do_execute_actions[2] the first n-1 times that an skb
> needs to be transmitted it is cloned first and the final time the original
> skb is used.
>
> In the case that there is only one action, which is the normal case, then
> the original skb will be used. But in the case of mirroring the cloning
> comes into effect. And in my case the cloned skb seems to go to the (slow)
> eth1 interface while the original skb goes to the (fast) dummy0 interface
> that I set up to be a mirror. The result is that dummy0 "paces" the flow,
> and its a cracking pace at that.
>
> As an experiment I hacked do_execute_actions() to use the original skb
> for the first action instead of the last one. In my case the result was
> that eth1 "paces" the flow, and things work reasonably nicely.
>
> Well, sort of. Things work well for non-GSO skbs but extremely poorly for
> GSO skbs where only 3 (yes 3, not 3%) end up at the remote host running
> netserv. I'm unsure why, but I digress.
>
> It seems to me that my hack illustrates the point that the flow ends up
> being "paced" by one interface. However I think that what would be
> desirable is that the flow is "paced" by the slowest link. Unfortunately
> I'm unsure how to achieve that.
>
Hi Simon !
"pacing" is done because skb is attached to a socket, and a socket has a
limited (but configurable) sndbuf. sk->sk_wmem_alloc is the current sum
of all truesize skbs in flight.
When you enter something that :
1) Get a clone of the skb, queue the clone to device X
2) queue the original skb to device Y
Then : Socket sndbuf is not affected at all by device X queue.
This is speed on device Y that matters.
You want to get servo control on both X and Y
You could try to
1) Get a clone of skb
Attach it to socket too (so that socket get a feedback of final
orphaning for the clone) with skb_set_owner_w()
queue the clone to device X
Unfortunatly, stacked skb->destructor() makes this possible only for
known destructor (aka sock_wfree())
> One idea that I had was to skb_get() the original skb each time it is
> cloned - that is easy enough. But unfortunately it seems to me that
> approach would require some sort of callback mechanism in kfree_skb() so
> that the cloned skbs can kfree_skb() the original skb.
>
> Ideas would be greatly appreciated.
>
> [1] http://openvswitch.org/pipermail/dev_openvswitch.org/2010-October/003806.html
> [2] http://openvswitch.org/cgi-bin/gitweb.cgi?p=openvswitch;a=blob;f=datapath/actions.c;h=5e16143ca402f7da0ee8fc18ee5eb16c3b7598e6;hb=HEAD
> --
^ permalink raw reply
* Flow Control and Port Mirroring Revisited
From: Simon Horman @ 2011-01-06 9:33 UTC (permalink / raw)
To: Rusty Russell
Cc: virtualization, Jesse Gross, dev, virtualization, netdev, kvm,
Michael S. Tsirkin
Hi,
Back in October I reported that I noticed a problem whereby flow control
breaks down when openvswitch is configured to mirror a port[1].
I have (finally) looked into this further and the problem appears to relate
to cloning of skbs, as Jesse Gross originally suspected.
More specifically, in do_execute_actions[2] the first n-1 times that an skb
needs to be transmitted it is cloned first and the final time the original
skb is used.
In the case that there is only one action, which is the normal case, then
the original skb will be used. But in the case of mirroring the cloning
comes into effect. And in my case the cloned skb seems to go to the (slow)
eth1 interface while the original skb goes to the (fast) dummy0 interface
that I set up to be a mirror. The result is that dummy0 "paces" the flow,
and its a cracking pace at that.
As an experiment I hacked do_execute_actions() to use the original skb
for the first action instead of the last one. In my case the result was
that eth1 "paces" the flow, and things work reasonably nicely.
Well, sort of. Things work well for non-GSO skbs but extremely poorly for
GSO skbs where only 3 (yes 3, not 3%) end up at the remote host running
netserv. I'm unsure why, but I digress.
It seems to me that my hack illustrates the point that the flow ends up
being "paced" by one interface. However I think that what would be
desirable is that the flow is "paced" by the slowest link. Unfortunately
I'm unsure how to achieve that.
One idea that I had was to skb_get() the original skb each time it is
cloned - that is easy enough. But unfortunately it seems to me that
approach would require some sort of callback mechanism in kfree_skb() so
that the cloned skbs can kfree_skb() the original skb.
Ideas would be greatly appreciated.
[1] http://openvswitch.org/pipermail/dev_openvswitch.org/2010-October/003806.html
[2] http://openvswitch.org/cgi-bin/gitweb.cgi?p=openvswitch;a=blob;f=datapath/actions.c;h=5e16143ca402f7da0ee8fc18ee5eb16c3b7598e6;hb=HEAD
^ permalink raw reply
* Re: [PATCH] virtio-pci: add softlinks between virtio and pci
From: Gleb Natapov @ 2011-01-06 8:51 UTC (permalink / raw)
To: Michael S. Tsirkin
Cc: Anthony Liguori, Jamie Lokier, Thomas Weber, linux-kernel,
virtualization
In-Reply-To: <20110105191711.GA27489@redhat.com>
On Wed, Jan 05, 2011 at 09:17:11PM +0200, Michael S. Tsirkin wrote:
> We sometimes need to map between the virtio device and
> the given pci device. One such use is OS installer that
> gets the boot pci device from BIOS and needs to
> find the relevant block device. Since it can't,
> installation fails.
>
> Supply softlinks between these to make it possible.
>
> Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
> ---
>
> Gleb, could you please ack that this patch below
> will be enough to fix the installer issue that
> you see?
>
ACK. With this patch given PCI address from EDD I can find device node
by looking at /sys/devices/pci0000:00/0000:edd_bdf/virtio_device/block/
> drivers/virtio/virtio_pci.c | 18 +++++++++++++++++-
> 1 files changed, 17 insertions(+), 1 deletions(-)
>
> diff --git a/drivers/virtio/virtio_pci.c b/drivers/virtio/virtio_pci.c
> index ef8d9d5..06eb2f8 100644
> --- a/drivers/virtio/virtio_pci.c
> +++ b/drivers/virtio/virtio_pci.c
> @@ -25,6 +25,7 @@
> #include <linux/virtio_pci.h>
> #include <linux/highmem.h>
> #include <linux/spinlock.h>
> +#include <linux/sysfs.h>
>
> MODULE_AUTHOR("Anthony Liguori <aliguori@us.ibm.com>");
> MODULE_DESCRIPTION("virtio-pci");
> @@ -667,8 +668,21 @@ static int __devinit virtio_pci_probe(struct pci_dev *pci_dev,
> if (err)
> goto out_set_drvdata;
>
> - return 0;
> + err = sysfs_create_link(&pci_dev->dev.kobj, &vp_dev->vdev.dev.kobj,
> + "virtio_device");
> + if (err)
> + goto out_register_device;
> +
> + err = sysfs_create_link(&vp_dev->vdev.dev.kobj, &pci_dev->dev.kobj,
> + "bus_device");
> + if (err)
> + goto out_create_link;
>
> + return 0;
> +out_create_link:
> + sysfs_remove_link(&pci_dev->dev.kobj, "virtio_device");
> +out_register_device:
> + unregister_virtio_device(&vp_dev->vdev);
> out_set_drvdata:
> pci_set_drvdata(pci_dev, NULL);
> pci_iounmap(pci_dev, vp_dev->ioaddr);
> @@ -685,6 +699,8 @@ static void __devexit virtio_pci_remove(struct pci_dev *pci_dev)
> {
> struct virtio_pci_device *vp_dev = pci_get_drvdata(pci_dev);
>
> + sysfs_remove_link(&vp_dev->vdev.dev.kobj, "bus_device");
> + sysfs_remove_link(&pci_dev->dev.kobj, "virtio_device");
> unregister_virtio_device(&vp_dev->vdev);
> }
>
> --
> 1.7.3.2.91.g446ac
--
Gleb.
^ permalink raw reply
* Re: [PATCH] virtio-pci: add softlinks between virtio and pci
From: Michael S. Tsirkin @ 2011-01-05 21:15 UTC (permalink / raw)
To: Anthony Liguori; +Cc: Jamie Lokier, Thomas Weber, linux-kernel, virtualization
In-Reply-To: <4D24CFC1.9020305@codemonkey.ws>
On Wed, Jan 05, 2011 at 02:08:33PM -0600, Anthony Liguori wrote:
> On 01/05/2011 02:05 PM, Michael S. Tsirkin wrote:
> >On Wed, Jan 05, 2011 at 01:28:34PM -0600, Anthony Liguori wrote:
> >>On 01/05/2011 01:17 PM, Michael S. Tsirkin wrote:
> >>>We sometimes need to map between the virtio device and
> >>>the given pci device. One such use is OS installer that
> >>>gets the boot pci device from BIOS and needs to
> >>>find the relevant block device. Since it can't,
> >>>installation fails.
> >>I have no objection to this patch but I'm a tad confused by the description.
> >>
> >>I assume you mean the installer is querying the boot device via
> >>int13 get driver parameters such that it returns the pci address of
> >>the device?
> >>
> >>Or is it querying geometry information and then trying to find the
> >>best match block device?
> >>
> >>If it's the former, I don't really understand the need for a
> >>backlink since the PCI address gives you a link to the block device.
> >>OTOH, if it's the later, it would make sense but then your
> >>description doesn't really make much sense.
> >>
> >>At any rate, a better commit message would be helpful in explaining
> >>the need for this.
> >>
> >>Regards,
> >>
> >>Anthony Liguori
> >OK just to clarify: we get pci address from BIOS
> >and need the virtio device to get at the linux device
> >(e.g. block) in the end. Thus the link from pci to virtio.
> >I also added a backlink since I thought it's handy.
> >
> >Does this answer the questions?
> >
> >Rusty rewrites my commit logs anyway, he has better style :)
>
> It helps. The real reason this is needed is because in a normal
> device, there is only one struct device whereas with virtio-pci, the
> virtio-pci device has a struct device and then the actual virtio
> device has another one.
I like how it works with e.g. net devices:
$ ls -l /sys/class/net/eth1
lrwxrwxrwx 1 root root 0 Jan 5 23:10 /sys/class/net/eth1 -> ../../devices/pci0000:00/0000:00:19.0/net/eth1
We maybe could have had
lrwxrwxrwx 1 root root 0 Jan 5 23:10 /sys/bus/virtio/virtio0 -> ../../devices/pci0000:00/0000:00:19.0/virtio/virtio0
This is pretty and would preserve the compatibility, but I am not sure how to implement this:
bus seems to want to have real kobjs behind it, not softlinks.
Ideas?
> There's probably a better way to handle
> this in sysfs making virtio-pci a proper bus with only a single
> device as a child or something like that.
I admit I'm just confused by all these buses.
>
> But the links are probably an easier solution.
>
> Regards,
>
> Anthony Liguori
>
> >>>Supply softlinks between these to make it possible.
> >>>
> >>>Signed-off-by: Michael S. Tsirkin<mst@redhat.com>
> >>>---
> >>>
> >>>Gleb, could you please ack that this patch below
> >>>will be enough to fix the installer issue that
> >>>you see?
> >>>
> >>> drivers/virtio/virtio_pci.c | 18 +++++++++++++++++-
> >>> 1 files changed, 17 insertions(+), 1 deletions(-)
> >>>
> >>>diff --git a/drivers/virtio/virtio_pci.c b/drivers/virtio/virtio_pci.c
> >>>index ef8d9d5..06eb2f8 100644
> >>>--- a/drivers/virtio/virtio_pci.c
> >>>+++ b/drivers/virtio/virtio_pci.c
> >>>@@ -25,6 +25,7 @@
> >>> #include<linux/virtio_pci.h>
> >>> #include<linux/highmem.h>
> >>> #include<linux/spinlock.h>
> >>>+#include<linux/sysfs.h>
> >>>
> >>> MODULE_AUTHOR("Anthony Liguori<aliguori@us.ibm.com>");
> >>> MODULE_DESCRIPTION("virtio-pci");
> >>>@@ -667,8 +668,21 @@ static int __devinit virtio_pci_probe(struct pci_dev *pci_dev,
> >>> if (err)
> >>> goto out_set_drvdata;
> >>>
> >>>- return 0;
> >>>+ err = sysfs_create_link(&pci_dev->dev.kobj,&vp_dev->vdev.dev.kobj,
> >>>+ "virtio_device");
> >>>+ if (err)
> >>>+ goto out_register_device;
> >>>+
> >>>+ err = sysfs_create_link(&vp_dev->vdev.dev.kobj,&pci_dev->dev.kobj,
> >>>+ "bus_device");
> >>>+ if (err)
> >>>+ goto out_create_link;
> >>>
> >>>+ return 0;
> >>>+out_create_link:
> >>>+ sysfs_remove_link(&pci_dev->dev.kobj, "virtio_device");
> >>>+out_register_device:
> >>>+ unregister_virtio_device(&vp_dev->vdev);
> >>> out_set_drvdata:
> >>> pci_set_drvdata(pci_dev, NULL);
> >>> pci_iounmap(pci_dev, vp_dev->ioaddr);
> >>>@@ -685,6 +699,8 @@ static void __devexit virtio_pci_remove(struct pci_dev *pci_dev)
> >>> {
> >>> struct virtio_pci_device *vp_dev = pci_get_drvdata(pci_dev);
> >>>
> >>>+ sysfs_remove_link(&vp_dev->vdev.dev.kobj, "bus_device");
> >>>+ sysfs_remove_link(&pci_dev->dev.kobj, "virtio_device");
> >>> unregister_virtio_device(&vp_dev->vdev);
> >>> }
> >>>
^ permalink raw reply
page: next (older) | prev (newer) | latest
- recent:[subjects (threaded)|topics (new)|topics (active)]
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox