sched/numa, mm: Use active_nodes nodemask to limit numa migrations Use the active_nodes nodemask to make smarter decisions on NUMA migrations. In order to maximize performance of workloads that do not fit in one NUMA node, we want to satisfy the following criteria: 1) keep private memory local to each thread 2) avoid excessive NUMA migration of pages 3) distribute shared memory across the active nodes, to maximize memory bandwidth available to the workload This patch accomplishes that by implementing the following policy for NUMA migrations: 1) always migrate on a private fault 2) never migrate to a node that is not in the set of active nodes for the numa_group 3) always migrate from a node outside of the set of active nodes, to a node that is in that set 4) within the set of active nodes in the numa_group, only migrate from a node with more NUMA page faults, to a node with fewer NUMA page faults, with a 25% margin to avoid ping-ponging This results in most pages of a workload ending up on the actively used nodes, with reduced ping-ponging of pages between those nodes. Signed-off-by: Rik van Riel <riel@redhat.com> Acked-by: Mel Gorman <mgorman@suse.de> Signed-off-by: Peter Zijlstra <peterz@infradead.org> Cc: Chegu Vinod <chegu_vinod@hp.com> Link: http://lkml.kernel.org/r/1390860228-21539-6-git-send-email-riel@redhat.com Signed-off-by: Ingo Molnar <mingo@kernel.org>

commit: 10f39042711ba21773763f267b4943a2c66c8bef [log] [tgz]
author: Rik van Riel <riel@redhat.com> Mon Jan 27 17:03:44 2014 -0500
committer: Ingo Molnar <mingo@kernel.org> Tue Jan 28 13:17:07 2014 +0100
tree: 394f8399c6f9b980f5673e1034b125b91844f662
parent: 20e07dea286a90f096a779706861472d296397c6 [diff] [blame]
diff --git a/include/linux/sched.h b/include/linux/sched.h
index 5fb0cfb..5ab3b89 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h

@@ -1589,6 +1589,8 @@
 extern pid_t task_numa_group_id(struct task_struct *p);
 extern void set_numabalancing_state(bool enabled);
 extern void task_numa_free(struct task_struct *p);
+extern bool should_numa_migrate_memory(struct task_struct *p, struct page *page,
+					int src_nid, int dst_cpu);
 #else
 static inline void task_numa_fault(int last_node, int node, int pages,
 				   int flags)
@@ -1604,6 +1606,11 @@
 static inline void task_numa_free(struct task_struct *p)
 {
 }
+static inline bool should_numa_migrate_memory(struct task_struct *p,
+				struct page *page, int src_nid, int dst_cpu)
+{
+	return true;
+}
 #endif
 
 static inline struct pid *task_pid(struct task_struct *task)
commit	10f39042711ba21773763f267b4943a2c66c8bef	[log] [tgz]
author	Rik van Riel <riel@redhat.com>	Mon Jan 27 17:03:44 2014 -0500
committer	Ingo Molnar <mingo@kernel.org>	Tue Jan 28 13:17:07 2014 +0100
tree	394f8399c6f9b980f5673e1034b125b91844f662
parent	20e07dea286a90f096a779706861472d296397c6 [diff] [blame]