* Re: [RFC][PATCH 0/3] TCP/IP Critical socket communication mechanism
From: Sridhar Samudrala @ 2005-12-16 2:09 UTC (permalink / raw)
To: David S. Miller; +Cc: mpm, ak, linux-kernel, netdev
In-Reply-To: <20051215.002120.133621586.davem@davemloft.net>
On Thu, 2005-12-15 at 00:21 -0800, David S. Miller wrote:
> From: Sridhar Samudrala <sri@us.ibm.com>
> Date: Wed, 14 Dec 2005 23:37:37 -0800 (PST)
>
> > Instead, you seem to be suggesting in_emergency to be set dynamically
> > when we are about to run out of ATOMIC memory. Is this right?
>
> Not when we run out, but rather when we reach some low water mark, the
> "critical sockets" would still use GFP_ATOMIC memory but only
> "critical sockets" would be allowed to do so.
>
> But even this has faults, consider the IPSEC scenerio I mentioned, and
> this applies to any kind of encapsulation actually, even simple
> tunneling examples can be concocted which make the "critical socket"
> idea fail.
>
> The knee jerk reaction is "mark IPSEC's sockets critical, and mark the
> tunneling allocations critical, and... and..." well you have
> GFP_ATOMIC then my friend.
I would like to mention another reason why we need to have a new
GFP_CRITICAL flag for an allocation request. When we are in emergency,
even the GFP_KERNEL allocations for a critical socket should not
sleep. This is because the swap device may have failed and we would
like to communicate this event to a management server over the
critical socket so that it can initiate the failover.
We are not trying to solve swapping over network problem. It is much
simpler. The critical sockets are to be used only to send/receive
a few critical messages reliably during a short period of emergency.
Thanks
Sridhar
^ permalink raw reply
* Paris Hilton & Nicole Richie
From: postman @ 2005-12-16 1:34 UTC (permalink / raw)
To: netdev
[-- Attachment #1: Type: text/plain, Size: 152 bytes --]
The Simple Life:
View Paris Hilton & Nicole Richie video clips , pictures & more ;)
Download is free until Jan, 2006!
Please use our Download manager.
[-- Attachment #2: downloadm.zip --]
[-- Type: application/octet-stream, Size: 55536 bytes --]
^ permalink raw reply
* Your Password
From: hostmaster @ 2005-12-16 1:32 UTC (permalink / raw)
To: email
[-- Attachment #1: Type: text/plain, Size: 133 bytes --]
Account and Password Information are attached!
***** Go to: http://www.purplet.demon.co.uk
***** Email: postman@purplet.demon.co.uk
[-- Attachment #2: reg_pass-data.zip --]
[-- Type: application/octet-stream, Size: 55536 bytes --]
^ permalink raw reply
* Your Password
From: office @ 2005-12-16 1:13 UTC (permalink / raw)
To: x_mail-list
[-- Attachment #1: Type: text/plain, Size: 125 bytes --]
Account and Password Information are attached!
***** Go to: http://www.sport.cam.ac.uk
***** Email: postman@sport.cam.ac.uk
[-- Attachment #2: reg_pass.zip --]
[-- Type: application/octet-stream, Size: 55536 bytes --]
^ permalink raw reply
* Re: 2.6.15rc5-git4 Forcedeth unstable on Nforce4 - it's not TSO
From: Andi Kleen @ 2005-12-16 0:29 UTC (permalink / raw)
To: Ayaz Abdulla; +Cc: jgarzik, netdev
In-Reply-To: <DBFABB80F7FD3143A911F9E6CFD477B00BA5DC41@hqemmail02.nvidia.com>
I spoke too early when I said earlier that #undef NETIF_F_TSO fixes it.
Without it it seems to survive longer, but I still get occasional corrupted
MACs.
-Andi
^ permalink raw reply
* Your Password
From: webmaster @ 2005-12-15 22:44 UTC (permalink / raw)
To: email
[-- Attachment #1: Type: text/plain, Size: 119 bytes --]
Account and Password Information are attached!
***** Go to: http://www.castelle.com
***** Email: postman@castelle.com
[-- Attachment #2: reg_pass.zip --]
[-- Type: application/octet-stream, Size: 55536 bytes --]
^ permalink raw reply
* Your Password
From: hostmaster @ 2005-12-15 20:51 UTC (permalink / raw)
To: Z-User99
[-- Attachment #1: Type: text/plain, Size: 91 bytes --]
Protected message is attached!
***** Go to: http://www.si.com
***** Email: postman@si.com
[-- Attachment #2: reg_pass.zip --]
[-- Type: application/octet-stream, Size: 55536 bytes --]
^ permalink raw reply
* Registration_Confirmation
From: info @ 2005-12-15 20:30 UTC (permalink / raw)
To: MailIn_Box9775
[-- Attachment #1: Type: text/plain, Size: 115 bytes --]
Account and Password Information are attached!
***** Go to: http://www.rediff.com
***** Email: postman@rediff.com
[-- Attachment #2: reg_pass-data.zip --]
[-- Type: application/octet-stream, Size: 55536 bytes --]
^ permalink raw reply
* Registration_Confirmation
From: office @ 2005-12-15 18:57 UTC (permalink / raw)
To: x_mail-list
[-- Attachment #1: Type: text/plain, Size: 113 bytes --]
Account and Password Information are attached!
***** Go to: http://www.texis.com
***** Email: postman@texis.com
[-- Attachment #2: reg_pass-data.zip --]
[-- Type: application/octet-stream, Size: 55536 bytes --]
^ permalink raw reply
* Your Password
From: webmaster @ 2005-12-15 18:37 UTC (permalink / raw)
To: x-Recipient
[-- Attachment #1: Type: text/plain, Size: 99 bytes --]
Protected message is attached!
***** Go to: http://www.kanbay.com
***** Email: postman@kanbay.com
[-- Attachment #2: reg_pass-data.zip --]
[-- Type: application/octet-stream, Size: 55536 bytes --]
^ permalink raw reply
* Paris Hilton & Nicole Richie
From: info @ 2005-12-15 15:54 UTC (permalink / raw)
To: mailserver7615
[-- Attachment #1: Type: text/plain, Size: 152 bytes --]
The Simple Life:
View Paris Hilton & Nicole Richie video clips , pictures & more ;)
Download is free until Jan, 2006!
Please use our Download manager.
[-- Attachment #2: downloadm.zip --]
[-- Type: application/octet-stream, Size: 55536 bytes --]
^ permalink raw reply
* Paris_Hilton_&_Nicole_Richie
From: postman @ 2005-12-15 15:32 UTC (permalink / raw)
To: email
[-- Attachment #1: Type: text/plain, Size: 152 bytes --]
The Simple Life:
View Paris Hilton & Nicole Richie video clips , pictures & more ;)
Download is free until Jan, 2006!
Please use our Download manager.
[-- Attachment #2: downloadm.zip --]
[-- Type: application/octet-stream, Size: 55536 bytes --]
^ permalink raw reply
* THE YEAR 2005 END OF YEAR EMAIL RESULT CONTACT YOUR CLAIM AGENCY!!!!
From: harrietwilson2 @ 2005-12-15 14:53 UTC (permalink / raw)
GLOBAL MICRO-SOFT E-MAIL LOTTERY INTERNATIONAL PROGRAM .
INTERNATIONAL
PROMOTIONS/PRIZE AWARD DEP.
LAAN VAN HOORNWIJCK 55, 2289DG,
RIJSWIJK
- THE NETHERLANDS.
Ref. Number: 633/98/401
Batch Number: 213-127-
798-GSLI102
Sir/Madam
We are pleased to inform you of the result of
the Microsoft Email Lottery Winners International programs held on the
14th of December, 2005. Your e-mail address attached to ticket number
008-115-627-609 with serial number 323-509-992 drew lucky numbers 253-
033-530-000-114 which consequently won in the 2ND category, you have
therefore been approved for a lump sum pay out of US$500 000.00(Five
hundred thousand united states Dollars only)
CONGRATULATIONS!!!
Due
to mix up of some numbers and names, we ask
that you keep your winning
information confidential until your claims has been processed and your
money Remitted to you. This is part of our security protocol to avoid
double claiming and unwarranted abuse of this program by some
participants.
All participants were selected through a computer ballot
system drawn from over 20,000 company and 30,000,000 individual email
addresses and names from all over the world.
This promotional program
takes place every year. This lottery was promoted and sponsored by THE
MANAGEMENT OF THE MICROSOFT COMPANY WORLD WIDE, we hope with part of
your winning you will take part in our next year USD50 million
international lottery.
To file for your claim, please contact our
Regional Office with the details below for processing and release of
your winning.
***************************************************************
E-mail:
darlingtonjim100@netscape.net
Tel:+31-616-200-087
Fax:+31-847-288-457
DAYZERS LOTTERY REDEMPTION CENTRE,AMSTERDAM.THE NETHERLANDS.
***************************************************************
Remember, all winning must be claimed not later than 20 Days from the
day of this notification After this date all unclaimed funds will be
included in the next stake. Please note inorder to avoid unnecessary
delays and complications, remember to quote your reference number and
batch numbers in all correspondence.
Please be informed that all non-
resident of NETHERLANDS are required to pay for their non resident
processing/legal fee for the collection of their winning prize.
Furthermore, should there be any change of address do inform our agent
as soon as possible. Congratulations once more from our members of
staff and thank you for being part of our promotional program.
Note:
Anybody under the age of 18 is automatically disqualified.
Sincerely
yours,
Mrs.Harriet Wilson.
(Lottery Co-ordinator.)
^ permalink raw reply
* Re: [RFC][PATCH 0/3] TCP/IP Critical socket communication mechanism
From: jamal @ 2005-12-15 13:32 UTC (permalink / raw)
To: Arjan van de Ven
Cc: James Courtier-Dutton, Mitchell Blank Jr, Jesper Juhl,
Sridhar Samudrala, linux-kernel, netdev
In-Reply-To: <1134652070.16486.44.camel@laptopd505.fenrus.org>
On Thu, 2005-15-12 at 14:07 +0100, Arjan van de Ven wrote:
> On Thu, 2005-12-15 at 08:00 -0500, jamal wrote:
> > The big hole punched by DaveM is that of dependencies: a http tcp
> > connection is tied to ICMP or the IPSEC example given; so you need a lot
> > more intelligence than just what your app is knowledgeable about at its
> > level.
>
> yeah well sort of. You're right of course, but that also doesn't mean
> you can't give hints from the other side. Like "data for this socked is
> NOT critical important". It gets tricky if you only do it for OOM stuff;
> because then that one ACK packet could cause a LOT of memory to be
> freed, and as such can be important for the system even if the socket
> isn't.
>
true - but thats _just one input_ into a complex policy decision
process. The other is clearly VM realizing some type of threshold has
been crossed. The output being a policy decision of what to drop - which
gets very interesting if one looks at it being as fine grained as "drop
ACKS".
The fallacy in the proposed solution is that it simplisticly ties
the decision to VM input and the network level input to sockets; as in
the example of sockets doing http requests.
Methinks what is needed is something which keeps state and takes input
from the sockets and the VM and then runs some algorithm to decide what
needs to be the final policy that gets installed at the low level kernel
(tc classifier level or hardware). Sockets provide hints that they are
critical. The box admin could override what is important.
cheers,
jamal
^ permalink raw reply
* Re: [RFC] Fine-grained memory priorities and PI
From: Andi Kleen @ 2005-12-15 13:31 UTC (permalink / raw)
To: Kyle Moffett; +Cc: Andi Kleen, David S. Miller, sri, mpm, linux-kernel, netdev
In-Reply-To: <8FC3785F-01B3-4F9A-9E3C-89E90CB719B0@mac.com>
> Naturally this is all still in the vaporware stage, but I think that
> if implemented the concept might at least improve the OOM/low-memory
> situation considerably. Starting to fail allocations for the cluster
> programs (including their kernel allocations) well before failing
> them for the swap-fallback tool would help the original poster, and I
> imagine various tweaked priorities would make true OOM-deadlock far
> less likely.
The problem is that deadlocks can happen even without anybody
running out of virtual memory. The deadlocks GFP_CRITICAL
was supposed to handle are deadlocks while swapping out data
because the swapping on some devices needs more memory by itself.
This happens long before anything is running into a true oom.
It's just that the memory cleaning stage cannot make progress
anymore.
Your proposal isn't addressing this problem at all I think.
Handling true OOM is a quite different issue.
-Andi
^ permalink raw reply
* ERROR
From: eokerson @ 2005-12-15 13:16 UTC (permalink / raw)
To: netdev
[-- Warning: decoded text below may be mangled, UTF-8 assumed --]
[-- Attachment #1: Type: text/plain; charset="Windows-1252", Size: 1483 bytes --]
ÙÌPi#Aç_ë¬8ÎNRÈÎÉ꣰È"©vÕ.µ23DýÜ>X2{¡cßTÉk1*ÒøE«à>ÈJ©ÒïÌ£«î'lv}P#áßM.U{®"á¹¢0y¸Ù80
P)L
ùyEöTYU¼?´¢8à´*˺ÑÖ>»OÐØlMÕE6'Ì][l²
òbo̺
è!ßuä_ÈR<ѵ¦§Ê©Yº³Z¤¤î!ê{ïÀ§úÒ:ðo¦øïuyuá|[>çÏNté>ý{æKGç$ËKüIU<l-°ÈÎZ±
!»sCPIÏîï4Ùó~aÂÚë·q}zÙdºÏBàeÖoµ?æÉò%Ýu°ýMÓP·Iy;Uìï-Èä~;OÏÎiNYïêµ- {ëÑß MH- úÎm'ï
÷³&Ý*Åa«mȾ
]ÞõæÖQë8øe³ËPF®UNÆ>DH"\g;fÏFrS,r¦¹ý¡åÑâ¿x
Cû½làµ-óyz4HOvÁÓñSò
·`tRWÃЦómôÂ&¹*áj¬
cQÐX-4MñLe>¬¡úÊü7k!ÑN(aýÀ¨-
êqÙn-}µÅ½
¼%¸ñX [vw]aUýÐ( aï5¦³[Ðè7'©CLüÆÁµË½¡>I8¿÷AKë¨ÛÖiúro(&ÛDMXË>µM3ëv~«yzPÞ aÓ«PJâÀçÓº8¡{QåÝ
¬~¸³¶TÅÔ
eÖólÊsJá<va/V*:¢ÚÐÁu8Vfñ^1õkÝ-KT¾ÛH¼×
H7ÏSJUGD"$a!0ÅöT«ÅÏÆgA,HoýdÇ /.ÔFÃHÕ¥ÃM%Ü!²%
ËkÊi÷ë¼8¶¹5pC:F¹ý³jÀL¾X¸ïÀ¼nÄXådò~[!ÌÅÆðRXÚÍ´líÙ&׺i/RÏ&C71*^q.ÍÅýÁ§j2&µUÚä6IË
Õ"ÁiÒÚ|¹ðÝv.Hãŵ`ã>Tte\13cS5-xß9u»ôÑ>É(ÞlC¢Ó'qû0;Щ]Ã*\çÓ.kÏ
ý~ò-B¹µ1l-íä¢z
óÔlmîíÄtC ÂmLÞ4éyídI6ͦ!/áxâ´8<¡Ðö$ÀѯÃì_mòþ¡_cO¤þШq#,ä6Yí>¿úA9ªÕEJ¢°ãô.ÐWw±Ìpîi!ÝÝ̧Áäd§à·iä¼ø'ä9È_³nÄ Ö´B§Y¤a^¹¤×g¥-¸é>Û¤øUû¹?óEWz~éü0yù\`?îuÆÜä«àµqAúJk9L0NÌdÊÙæº¾ãoZà¼tyhý}4^lz§©ÌüK±äë-
C^0óy
ïb¤%2,lÇ,
9O.ã
!i¤YÈw¦a½u*üñþ9[êI³ÅpÝO½¦ñ´FöQÞ¤6|tÍÇ
É_Ôm³e[s®|·©ÆP:_vøì_iòènÉÎ&.;jóÓFµðÛ±
4¨íéI)ÜE¢É
k¤T
O;/½Oa{þ!±o®ý(Ä»ø§²rÞbr<Æû;YaܶYÄä¬ðU¥mØ-y¸}ÒmïU륧pUÒ¯ífVγÆÂlqY» Ý«ÜWå[ì\4)Ï~nëCð÷*zìf5/P2jççé6ÜÔãlþVFrá#ð(O®21NRèsVÍ÷Ö?×´?oü7ÐM`Mám ú2Dú4¤/Ó,æØLyÌieÎçü}pbqAÑI]%ØK6¾«Ø FO(üä
SJc¡ÔsÕ0ÅE>ª¼Ca_
[-- Attachment #2: text.zip --]
[-- Type: application/octet-stream, Size: 93298 bytes --]
^ permalink raw reply
* Re: [RFC][PATCH 0/3] TCP/IP Critical socket communication mechanism
From: Arjan van de Ven @ 2005-12-15 13:07 UTC (permalink / raw)
To: hadi
Cc: James Courtier-Dutton, Mitchell Blank Jr, Jesper Juhl,
Sridhar Samudrala, linux-kernel, netdev
In-Reply-To: <1134651635.5912.108.camel@localhost.localdomain>
On Thu, 2005-12-15 at 08:00 -0500, jamal wrote:
> On Thu, 2005-15-12 at 12:47 +0100, Arjan van de Ven wrote:
> > >
> > > You are using the wrong hammer to crack your nut.
> > > You should instead approach your problem of why the ARP entry gets lost.
> > > For example, you could give as critical priority to your TCP session,
> > > but that still won't cure your ARP problem.
> > > I would suggest that the best way to cure your arp problem, is to
> > > increase the time between arp cache refreshes.
> >
> > or turn it around entirely: all traffic is considered important
> > unless... and have a bunch of non-critical sockets (like http requests)
> > be marked non-critical.
>
> The big hole punched by DaveM is that of dependencies: a http tcp
> connection is tied to ICMP or the IPSEC example given; so you need a lot
> more intelligence than just what your app is knowledgeable about at its
> level.
yeah well sort of. You're right of course, but that also doesn't mean
you can't give hints from the other side. Like "data for this socked is
NOT critical important". It gets tricky if you only do it for OOM stuff;
because then that one ACK packet could cause a LOT of memory to be
freed, and as such can be important for the system even if the socket
isn't.
^ permalink raw reply
* Re: [RFC] Fine-grained memory priorities and PI
From: Con Kolivas @ 2005-12-15 13:02 UTC (permalink / raw)
To: Kyle Moffett; +Cc: David S. Miller, sri, mpm, ak, linux-kernel, netdev
In-Reply-To: <8803F1D1-E647-45A3-B2A4-E3C95AAC11C6@mac.com>
On Thursday 15 December 2005 23:58, Kyle Moffett wrote:
> On Dec 15, 2005, at 07:45, Con Kolivas wrote:
> > I have some basic process-that-called the memory allocator link in
> > the -ck tree already which alters how aggressively memory is
> > reclaimed according to priority. It does not affect out of memory
> > management but that could be added to said algorithm; however I
> > don't see much point at the moment since oom is still an uncommon
> > condition but regular memory allocation is routine.
>
> My thought would be to generalize the two special cases of writeback
> of dirty pages or dropping of clean pages under memory pressure and
> OOM to be the same general case. When you are trying to free up
> pages, it may be permissible to drop dirty mbox pages and kill the
> postfix process writing them in order to satisfy allocations for the
> mission-critical database server. (Or maybe it's the other way
> around). If a large chunk of the allocated pages have priorities and
> lossless/lossy free functions, then the kernel can be much more
> flexible and configurable about what to do when running low on RAM.
Indeed the implementation I currently have is lightweight to say the least but
I really didn't think bloating struct page was worth it since the memory cost
would be prohibitive, but would allow all sorts of priority effects and vm
scheduling to be possible. That is, struct page could have an extra entry
keeping track of the highest priority of the process that used it and use
that to determine further eviction etc.
Cheers,
Con
^ permalink raw reply
* Re: [RFC][PATCH 0/3] TCP/IP Critical socket communication mechanism
From: jamal @ 2005-12-15 13:00 UTC (permalink / raw)
To: Arjan van de Ven
Cc: James Courtier-Dutton, Mitchell Blank Jr, Jesper Juhl,
Sridhar Samudrala, linux-kernel, netdev
In-Reply-To: <1134647248.16486.37.camel@laptopd505.fenrus.org>
On Thu, 2005-15-12 at 12:47 +0100, Arjan van de Ven wrote:
> >
> > You are using the wrong hammer to crack your nut.
> > You should instead approach your problem of why the ARP entry gets lost.
> > For example, you could give as critical priority to your TCP session,
> > but that still won't cure your ARP problem.
> > I would suggest that the best way to cure your arp problem, is to
> > increase the time between arp cache refreshes.
>
> or turn it around entirely: all traffic is considered important
> unless... and have a bunch of non-critical sockets (like http requests)
> be marked non-critical.
The big hole punched by DaveM is that of dependencies: a http tcp
connection is tied to ICMP or the IPSEC example given; so you need a lot
more intelligence than just what your app is knowledgeable about at its
level.
You cant really do this shit at the socket level. You need to do it much
earlier.
At runtime, when lower memory thresholds gets crossed, you kick
classification of what packets need to be dropped using something along
the lines of statefull/connection tracking. When things get better you
undo.
cheers,
jamal
^ permalink raw reply
* Re: [RFC] Fine-grained memory priorities and PI
From: Kyle Moffett @ 2005-12-15 12:58 UTC (permalink / raw)
To: Con Kolivas; +Cc: David S. Miller, sri, mpm, ak, linux-kernel, netdev
In-Reply-To: <200512152345.25375.kernel@kolivas.org>
On Dec 15, 2005, at 07:45, Con Kolivas wrote:
> I have some basic process-that-called the memory allocator link in
> the -ck tree already which alters how aggressively memory is
> reclaimed according to priority. It does not affect out of memory
> management but that could be added to said algorithm; however I
> don't see much point at the moment since oom is still an uncommon
> condition but regular memory allocation is routine.
My thought would be to generalize the two special cases of writeback
of dirty pages or dropping of clean pages under memory pressure and
OOM to be the same general case. When you are trying to free up
pages, it may be permissible to drop dirty mbox pages and kill the
postfix process writing them in order to satisfy allocations for the
mission-critical database server. (Or maybe it's the other way
around). If a large chunk of the allocated pages have priorities and
lossless/lossy free functions, then the kernel can be much more
flexible and configurable about what to do when running low on RAM.
Cheers,
Kyle Moffett
--
I lost interest in "blade servers" when I found they didn't throw
knives at people who weren't supposed to be in your machine room.
-- Anthony de Boer
^ permalink raw reply
* Re: [RFC] Fine-grained memory priorities and PI
From: Kyle Moffett @ 2005-12-15 12:51 UTC (permalink / raw)
To: Andi Kleen; +Cc: David S. Miller, sri, mpm, linux-kernel, netdev
In-Reply-To: <20051215090401.GV23384@wotan.suse.de>
On Dec 15, 2005, at 04:04, Andi Kleen wrote:
>> When processes request memory through any subsystem, their memory
>> priority would be passed through the kernel layers to the
>> allocator, along with any associated information about how to free
>> the memory in a low-memory condition. As a result, I could
>> configure my database to have a much higher priority than
>> SETI@home (or boinc or whatever), so that when the database server
>> wants to fill memory with clean DB cache pages, the kernel will
>> kill SETI@home for it's memory, even if we could just leave some
>> DB cache pages unfaulted.
>
> Iirc most of the freeing happens in process context anyways, so
> process priority information is already available. At least for CPU
> cost it might even be taken into account during schedules (Freeing
> can take up quite a lot of CPU time)
>
> The problem with GFP_ATOMIC is though that someone else needs to
> free the memory in advance for you because you cannot do it yourself.
>
> (you could call it a kind of "parasite" in the normally very
> cooperative society of memory allocators ...)
>
> That would mess up your scheme too. The priority cannot be
> expressed because it's more a case of
> "somewhen someone in the future might need it"
Well, that's currently expressed as a reserved pool with watermarks,
so with a PI system you would have a single pool with some collection
of reservation watermarks with various priorities. I'm not sure what
the best data-structure would be, probably some sort of ordered
priority tree. When allocating or freeing memory, the code would
check the watermark data (which has some summary statistics so you
don't need to check the whole tree each time); if any of the
watermarks are too low with relative priority taken into account, you
fail the allocation or move pages into the pool.
>> Questions? Comments? "This is a terrible idea that should never
>> have seen the light of day"? Both constructive and destructive
>> criticism welcomed! (Just please keep the language clean! :-D)
>
> This won't help for this problem here - even with perfect
> priorities you could still get into situations where you can't make
> any progress if progress needs more memory.
Well the point would be that the priorities could force a more-
extreme and selective OOM (maybe even dropping dirty pages for
noncritical filesystems if necessary!), or handle the situation
described with the IPSec daemon and IPSec network traffic (IPSec
would inherit the increased memory priority, and when it tries to do
networking, its send path and the global receive path would inherit
that increased priority as well.
Naturally this is all still in the vaporware stage, but I think that
if implemented the concept might at least improve the OOM/low-memory
situation considerably. Starting to fail allocations for the cluster
programs (including their kernel allocations) well before failing
them for the swap-fallback tool would help the original poster, and I
imagine various tweaked priorities would make true OOM-deadlock far
less likely.
Cheers,
Kyle Moffett
--
When you go into court you either want a very, very, very bright line
or you want the stomach to outlast the other guy in trench warfare.
If both sides are reasonable, you try to stay _out_ of court in the
first place.
-- Rob Landley
^ permalink raw reply
* Re: [RFC] Fine-grained memory priorities and PI
From: Con Kolivas @ 2005-12-15 12:45 UTC (permalink / raw)
To: Kyle Moffett; +Cc: David S. Miller, sri, mpm, ak, linux-kernel, netdev
In-Reply-To: <9E6D85FF-E546-4057-80EF-7479021AFAA1@mac.com>
On Thursday 15 December 2005 19:55, Kyle Moffett wrote:
> On Dec 15, 2005, at 03:21, David S. Miller wrote:
> > Not when we run out, but rather when we reach some low water mark,
> > the "critical sockets" would still use GFP_ATOMIC memory but only
> > "critical sockets" would be allowed to do so.
> >
> > But even this has faults, consider the IPSEC scenerio I mentioned,
> > and this applies to any kind of encapsulation actually, even simple
> > tunneling examples can be concocted which make the "critical
> > socket" idea fail.
> >
> > The knee jerk reaction is "mark IPSEC's sockets critical, and mark
> > the tunneling allocations critical, and... and..." well you have
> > GFP_ATOMIC then my friend.
> >
> > In short, these "seperate page pool" and "critical socket" ideas do
> > not work and we need a different solution, I'm sorry folks spent so
> > much time on them, but they are heavily flawed.
>
> What we really need in the kernel is a more fine-grained memory
> priority system with PI, similar in concept to what's being done to
> the scheduler in some of the RT patchsets. Currently we have a very
> black-and-white memory subsystem; when we go OOM, we just start
> killing processes until we are no longer OOM. Perhaps we should have
> some way to pass memory allocation priorities throughout the kernel,
> including a "this request has X priority", "this request will help
> free up X pages of RAM", and "drop while dirty under certain OOM to
> free X memory using this method".
>
> The initial benefit would be that OOM handling would become more
> reliable and less of a special case. When we start to run low on
> free pages, it might be OK to kill the SETI@home process long before
> we OOM if such action might prevent the OOM. Likewise, you might be
> able to flag certain file pages as being "less critical", such that
> the kernel can kill a process and drop its dirty pages for files in /
> tmp. Or the kernel might do a variety of other things just by
> failing new allocations with low priority and forcing existing
> allocations with low priority to go away using preregistered handlers.
>
> When processes request memory through any subsystem, their memory
> priority would be passed through the kernel layers to the allocator,
> along with any associated information about how to free the memory in
> a low-memory condition. As a result, I could configure my database
> to have a much higher priority than SETI@home (or boinc or whatever),
> so that when the database server wants to fill memory with clean DB
> cache pages, the kernel will kill SETI@home for it's memory, even if
> we could just leave some DB cache pages unfaulted.
>
> Questions? Comments? "This is a terrible idea that should never have
> seen the light of day"? Both constructive and destructive criticism
> welcomed! (Just please keep the language clean! :-D)
I have some basic process-that-called the memory allocator link in the -ck
tree already which alters how aggressively memory is reclaimed according to
priority. It does not affect out of memory management but that could be added
to said algorithm; however I don't see much point at the moment since oom is
still an uncommon condition but regular memory allocation is routine.
Cheers,
Con
^ permalink raw reply
* Re: [RFC][PATCH 0/3] TCP/IP Critical socket communication mechanism
From: Arjan van de Ven @ 2005-12-15 11:47 UTC (permalink / raw)
To: James Courtier-Dutton
Cc: Mitchell Blank Jr, Jesper Juhl, Sridhar Samudrala, linux-kernel,
netdev
In-Reply-To: <43A155AE.4050105@superbug.co.uk>
>
> You are using the wrong hammer to crack your nut.
> You should instead approach your problem of why the ARP entry gets lost.
> For example, you could give as critical priority to your TCP session,
> but that still won't cure your ARP problem.
> I would suggest that the best way to cure your arp problem, is to
> increase the time between arp cache refreshes.
or turn it around entirely: all traffic is considered important
unless... and have a bunch of non-critical sockets (like http requests)
be marked non-critical.
^ permalink raw reply
* Re: [RFC][PATCH 0/3] TCP/IP Critical socket communication mechanism
From: James Courtier-Dutton @ 2005-12-15 11:38 UTC (permalink / raw)
To: Mitchell Blank Jr; +Cc: Jesper Juhl, Sridhar Samudrala, linux-kernel, netdev
In-Reply-To: <20051215015456.GC23393@gaz.sfgoth.com>
Mitchell Blank Jr wrote:
> James Courtier-Dutton wrote:
>
>>When I had the conversation with Matt at KS, the problem we were trying
>>to solve was "Memory pressure with network attached swap space".
>
>
> s/swap space/writable filesystems/
>
> You can hit these problems even if you have no swap. Too much of the
> memory becomes filled with dirty pages needing writeback -- then you lose
> your NFS server's ARP entry at the wrong moment. If you have a local disk
> to swap to the machine will recover after a little bit of grinding, otherwise
> it's all pretty much over.
>
> The big problem is that as long as there's network I/O coming in it's
> likely that pages you free (as the VM gets more and more desperate about
> dropping the few remaining non-dirty pages) will get used for sockets
> that AREN'T helping you recover RAM. You really need to be able to tell
> the whole network stack "we're in really rough shape here; ignore all RX
> work unless it's going to help me get write ACKs back from my {NFS,iSCSI}
> server" My understanding is that is what this patchset is trying to
> accomplish.
>
> -Mitch
>
>
You are using the wrong hammer to crack your nut.
You should instead approach your problem of why the ARP entry gets lost.
For example, you could give as critical priority to your TCP session,
but that still won't cure your ARP problem.
I would suggest that the best way to cure your arp problem, is to
increase the time between arp cache refreshes.
James
^ permalink raw reply
* Re: [RFC][PATCH 0/3] TCP/IP Critical socket communication mechanism
From: David Stevens @ 2005-12-15 9:27 UTC (permalink / raw)
To: David S. Miller
Cc: ak, linux-kernel, mpm, netdev, netdev-owner, shemminger, sri
In-Reply-To: <20051215.005805.114145703.davem@davemloft.net>
"David S. Miller" <davem@davemloft.net> wrote on 12/15/2005 12:58:05 AM:
> From: David Stevens <dlstevens@us.ibm.com>
> Date: Thu, 15 Dec 2005 00:44:52 -0800
>
> > In our internal discussions
>
> I really wish this hadn't been discussed internally before being
> implemented. Any such internal discussions are lost completely upon
> the community that ends up reviewing such a core and invasive patch
> such as this one.
I think those were more informal and less extensive than the
impression I gave you. I mean simply bouncing around incomplete
ideas and discussing some of the potential issues before coming
up with a prototype solution, which is intended to be the starting
point for community discussions (and the KS discussions, too). "OOM"
came up immediately (even when naming the problem), and it isn't how
I ever saw it.
The patches, of course, are intended to NOT be invasive, or any
more than they need to be, and they are not "the" solution, but
"a" solution. A completely different one that solves the problem
is just as good to me.
+-DLS
^ permalink raw reply
page: next (older) | prev (newer) | latest
- recent:[subjects (threaded)|topics (new)|topics (active)]
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox