Same issue with 4.4.0-31-generic, stack trace: [443830.036000] divide error: 0000 [#1] SMP [443830.036583] Modules linked in: nf_conntrack_netlink xt_multiport xt_CT xt_mac xt_physdev xt_set ip_set_hash_net ip_set nfnetlink vhost_net vhost macvtap macvlan xt_REDIRECT nf_nat_redirect xt_mark vport_vxlan xt_CHECKSUM ip6table_raw nf_conntrack_ipv6 ip6table_mangle xt_connmark xt_comment iptable_raw iptable_mangle dccp_diag dccp tcp_diag udp_diag inet_diag unix_diag ebtable_filter ebtables ip6table_filter ip6_tables ip_vs openvswitch nf_defrag_ipv6 xt_nat xt_tcpudp veth ipt_MASQUERADE nf_nat_masquerade_ipv4 iptable_nat nf_conntrack_ipv4 nf_defrag_ipv4 nf_nat_ipv4 xt_addrtype iptable_filter ip_tables xt_conntrack x_tables nf_nat nf_conntrack br_netfilter bridge aufs mpt3sas raid_class scsi_transport_sas mptctl mptbase binfmt_misc bonding xfs nls_iso8859_1 intel_rapl ipmi_ssif joydev input_leds x86_pkg_temp_thermal [443830.039865] intel_powerclamp coretemp sb_edac mei_me mei edac_core lpc_ich ioatdma ipmi_si ipmi_msghandler shpchp 8250_fintek acpi_power_meter mac_hid kvm_intel acpi_pad kvm irqbypass ib_iser rdma_cm iw_cm ib_cm ib_sa ib_mad ib_core ib_addr iscsi_tcp libiscsi_tcp libiscsi scsi_transport_iscsi 8021q garp mrp stp llc sunrpc autofs4 btrfs raid10 raid456 async_raid6_recov async_memcpy async_pq async_xor async_tx xor raid6_pq libcrc32c raid1 raid0 multipath linear ses enclosure hid_generic crct10dif_pclmul ast crc32_pclmul i2c_algo_bit ttm ixgbe drm_kms_helper aesni_intel dca aes_x86_64 vxlan lrw ip6_udp_tunnel gf128mul syscopyarea glue_helper udp_tunnel ablk_helper sysfillrect cryptd usbhid sysimgblt fb_sys_fops hid ahci ptp drm libahci megaraid_sas pps_core mdio wmi fjes [443830.045336] CPU: 10 PID: 13866 Comm: ceph-osd Not tainted 4.4.0-31-generic #50-Ubuntu [443830.046219] Hardware name: Supermicro PIO-628U-TR4T+-ST031/X10DRU-i+, BIOS 2.0 12/17/2015 [443830.047112] task: ffff881fc1260dc0 ti: ffff881fc126c000 task.ti: ffff881fc126c000 [443830.048073] RIP: 0010:[<ffffffff810b5adc>] [<ffffffff810b5adc>] task_numa_find_cpu+0x23c/0x710 [443830.049023] RSP: 0000:ffff881fc126fbd8 EFLAGS: 00010206 [443830.050042] RAX: 0000000000000000 RBX: ffff881fc126fc78 RCX: 0000000000000000 [443830.051023] RDX: 0000000000000000 RSI: 0000000000000001 RDI: ffff881fef0c0800 [443830.052067] RBP: ffff881fc126fc40 R08: 00000001069be1a6 R09: 0000000000000015 [443830.053080] R10: 00000000000003d2 R11: 0000000000000df4 R12: ffff883b2e116e00 [443830.054174] R13: 000000000000000c R14: 0000000000000000 R15: fffffffffffffca6 [443830.055146] FS: 00007fe71a644700(0000) GS:ffff881fffa80000(0000) knlGS:0000000000000000 [443830.056197] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [443830.057206] CR2: 000000000aea4980 CR3: 0000001feed55000 CR4: 00000000001426e0 [443830.058296] Stack: [443830.059327] ffff881fc126fbf8 ffffffff813f0f4f 0000000000000100 ffff881fc1260dc0 [443830.060462] 0000000000000063 0000000000000129 0000000000016d00 0000000000000063 [443830.061607] ffff881fc1260dc0 ffff881fc126fc78 000000000000015f 00000000000001c2 [443830.062721] Call Trace: [443830.063831] [<ffffffff813f0f4f>] ? cpumask_next_and+0x2f/0x40 [443830.064965] [<ffffffff810b63ee>] task_numa_migrate+0x43e/0x9b0 [443830.066136] [<ffffffff8105a4fe>] ? physflat_send_IPI_mask+0xe/0x10 [443830.067227] [<ffffffff81037e99>] ? sched_clock+0x9/0x10 [443830.068314] [<ffffffff810b69d9>] numa_migrate_preferred+0x79/0x80 [443830.069490] [<ffffffff810baff4>] task_numa_fault+0x7f4/0xd40 [443830.070987] [<ffffffff810ba665>] ? should_numa_migrate_memory+0x55/0x130 [443830.072113] [<ffffffff811bffa0>] handle_mm_fault+0xbc0/0x1820 [443830.073248] [<ffffffff81102f53>] ? do_futex+0xd3/0x540 [443830.074395] [<ffffffff8106b537>] __do_page_fault+0x197/0x400 [443830.075538] [<ffffffff8106b7c2>] do_page_fault+0x22/0x30 [443830.076708] [<ffffffff8182fcb8>] page_fault+0x28/0x30 [443830.077862] Code: 55 b0 4c 89 f7 e8 25 c8 ff ff 48 8b 55 b0 49 8b 4e 78 48 8b 82 d8 01 00 00 48 83 c1 01 31 d2 49 0f af 86 b0 00 00 00 4c 8b 73 78 <48> f7 f1 48 8b 4b 20 49 89 c0 48 29 c1 48 8b 45 d0 4c 03 43 48 [443830.080280] RIP [<ffffffff810b5adc>] task_numa_find_cpu+0x23c/0x710 [443830.081470] RSP <ffff881fc126fbd8> [443830.086523] ---[ end trace 0f566374d1589a3d ]---
-- You received this bug notification because you are a member of Kernel Packages, which is subscribed to linux in Ubuntu. https://bugs.launchpad.net/bugs/1568729 Title: divide error: 0000 [#1] SMP in task_numa_migrate - handle_mm_fault Status in linux package in Ubuntu: In Progress Status in linux source package in Xenial: In Progress Bug description: While running qemu 2.5 on a trusty host running 4.4.0-15.31~14.04.1 the host system has crashed (load > 200) 3 times in the last 3 days. Always with this stack trace: Apr 9 19:01:09 cnode9.0 kernel: [197071.195577] divide error: 0000 [#1] SMP Apr 9 19:01:09 cnode9.0 kernel: [197071.195633] Modules linked in: vhost_net vhost macvtap macvlan arc4 md4 nls_utf8 ci fs nfnetlink_queue nfnetlink xt_CHECKSUM xt_nat iptable_nat nf_nat_ipv4 xt_NFQUEUE xt_CLASSIFY ip6table_mangle sch_sfq sch_htb veth dccp_diag dccp tcp_diag udp_diag inet_diag unix_diag af_packet_diag netlink_diag ebtable_filter ebtables nf_conntrack_ipv6 nf_defrag_ipv6 ip6table_fil ter ip6_tables iptable_mangle xt_CT iptable_raw xt_tcpudp nf_conntrack_ipv4 nf_defrag_ipv4 xt_conntrack iptable_filter ip_tables x_tables dum my bridge stp llc ipmi_ssif ipmi_devintf intel_rapl x86_pkg_temp_thermal intel_powerclamp coretemp kvm_intel kvm dcdbas irqbypass crct10dif_p clmul crc32_pclmul aesni_intel aes_x86_64 lrw gf128mul glue_helper ablk_helper cryptd joydev input_leds nf_nat_ftp sb_edac nf_conntrack_ftp e dac_core cdc_ether nf_nat_pptp usbnet nf_conntrack_pptp mii nf_nat_proto_gre lpc_ich nf_nat_sip ioatdma nf_nat nf_conntrack_sip nfsd ipmi_si 8250_fintek nf_conntrack_proto_gre ipmi_msghandler acpi_pad wmi shpchp nf_conntrack acpi_power_meter mac_hid auth_rpcgss nfs_acl bonding nfs lp lockd parport grace sunrpc fscache tcp_htcp xfs btrfs hid_generic usbhid hid raid10 raid456 async_raid6_recov async_memcpy async_pq async_ xor async_tx xor ixgbe raid6_pq libcrc32c igb vxlan raid1 i2c_algo_bit ip6_udp_tunnel dca udp_tunnel ahci raid0 ptp libahci megaraid_sas mult ipath pps_core mdio linear fjes Apr 9 19:01:09 cnode9.0 kernel: [197071.197014] CPU: 13 PID: 3147726 Comm: ceph-osd Not tainted 4.4.0-15-generic #31~14 .04.1-Ubuntu Apr 9 19:01:09 cnode9.0 kernel: [197071.197085] Hardware name: Dell Inc. PowerEdge R720/0XH7F2, BIOS 2.5.2 01/28/2015 Apr 9 19:01:09 cnode9.0 kernel: [197071.197154] task: ffff88252be1ee00 ti: ffff8824fc0d4000 task.ti: ffff8824fc0d4000 Apr 9 19:01:09 cnode9.0 kernel: [197071.197221] RIP: 0010:[<ffffffff810afec8>] [<ffffffff810afec8>] task_numa_find_cpu+0x238/0x700 Apr 9 19:01:09 cnode9.0 kernel: [197071.197300] RSP: 0000:ffff8824fc0d7ba8 EFLAGS: 00010257 Apr 9 19:01:09 cnode9.0 kernel: [197071.197340] RAX: 0000000000000000 RBX: ffff8824fc0d7c48 RCX: 0000000000000000 Apr 9 19:01:09 cnode9.0 kernel: [197071.197406] RDX: 0000000000000000 RSI: ffff88479f180000 RDI: ffff884782a47600 Apr 9 19:01:09 cnode9.0 kernel: [197071.197473] RBP: ffff8824fc0d7c10 R08: 0000000102eea157 R09: 00000000000001a8 Apr 9 19:01:09 cnode9.0 kernel: [197071.197540] R10: 000000000002404b R11: 000000000000023f R12: ffff882380930000 Apr 9 19:01:09 cnode9.0 kernel: [197071.197606] R13: 0000000000000008 R14: 000000000000008c R15: 0000000000000124 Apr 9 19:01:09 cnode9.0 kernel: [197071.197673] FS: 00007f19aab5b700(0000) GS:ffff88479f180000(0000) knlGS:0000000000000000 Apr 9 19:01:09 cnode9.0 kernel: [197071.197741] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 Apr 9 19:01:09 cnode9.0 kernel: [197071.197782] CR2: 0000000025469600 CR3: 00000023846bc000 CR4: 00000000000426e0 Apr 9 19:01:09 cnode9.0 kernel: [197071.197848] Stack: Apr 9 19:01:09 cnode9.0 kernel: [197071.197880] ffffffff817425fb ffff8829af3e9e00 00000000000000f6 ffff88252be1ee00 Apr 9 19:01:09 cnode9.0 kernel: [197071.197965] 000000000000008d 0000000000000225 0000000000016d40 000000000000008d Apr 9 19:01:09 cnode9.0 kernel: [197071.198047] ffff88252be1ee00 00000000000001ad ffff8824fc0d7c48 00000000000000e1 Apr 9 19:01:09 cnode9.0 kernel: [197071.198132] Call Trace: Apr 9 19:01:09 cnode9.0 kernel: [197071.198172] [<ffffffff817425fb>] ? tcp_schedule_loss_probe+0x12b/0x1b0 Apr 9 19:01:09 cnode9.0 kernel: [197071.198219] [<ffffffff810b0830>] task_numa_migrate+0x4a0/0x930 Apr 9 19:01:09 cnode9.0 kernel: [197071.198264] [<ffffffff816d2957>] ? release_sock+0x117/0x160 Apr 9 19:01:09 cnode9.0 kernel: [197071.198306] [<ffffffff810b0d39>] numa_migrate_preferred+0x79/0x80 Apr 9 19:01:09 cnode9.0 kernel: [197071.198350] [<ffffffff810b557d>] task_numa_fault+0x91d/0xcc0 Apr 9 19:01:09 cnode9.0 kernel: [197071.198395] [<ffffffff811d35ae>] ? mpol_misplaced+0x14e/0x190 Apr 9 19:01:09 cnode9.0 kernel: [197071.198439] [<ffffffff811b06b8>] handle_pte_fault+0x5a8/0x14c0 Apr 9 19:01:09 cnode9.0 kernel: [197071.198485] [<ffffffff810f8531>] ? futex_wake+0x81/0x150 Apr 9 19:01:09 cnode9.0 kernel: [197071.198526] [<ffffffff810b0de4>] ? set_next_entity+0xa4/0x700 Apr 9 19:01:09 cnode9.0 kernel: [197071.198569] [<ffffffff810fab44>] ? do_futex+0xf4/0x4d0 Apr 9 19:01:09 cnode9.0 kernel: [197071.198610] [<ffffffff811b2440>] handle_mm_fault+0x250/0x540 Apr 9 19:01:09 cnode9.0 kernel: [197071.198654] [<ffffffff81067d19>] __do_page_fault+0x199/0x430 Apr 9 19:01:09 cnode9.0 kernel: [197071.198696] [<ffffffff81067fd2>] do_page_fault+0x22/0x30 Apr 9 19:01:09 cnode9.0 kernel: [197071.198740] [<ffffffff817ef878>] page_fault+0x28/0x30 Apr 9 19:01:09 cnode9.0 kernel: [197071.198775] Code: 4d b0 4c 89 f7 e8 29 d5 ff ff 48 8b 4d b0 49 8b 86 b0 00 00 00 31 d2 48 0f af 81 d8 01 00 00 49 8b 4e 78 4c 8b 73 78 48 83 c1 01 <48> f7 f1 48 8b 4b 20 49 89 c1 48 29 c1 4c 03 4b 48 4c 39 7d d0 Apr 9 19:01:09 cnode9.0 kernel: [197071.199217] RIP [<ffffffff810afec8>] task_numa_find_cpu+0x238/0x700 Apr 9 19:01:09 cnode9.0 kernel: [197071.199264] RSP <ffff8824fc0d7ba8> Apr 9 19:01:09 cnode9.0 kernel: [197071.199900] ---[ end trace e938a840610a79f7 ]--- This is appears to be the same bug as reported upstream in http://lkml.iu.edu/hypermail/linux/kernel/1603.2/01659.html According to this thread the issue is: 27: 48 83 c1 01 add $0x1,%rcx 2b:* 48 f7 f1 div %rcx <-- trapping instruction This suggests the CONFIG_FAIR_GROUP_SCHED version of task_h_load: update_cfs_rq_h_load(cfs_rq); return div64_ul(p->se.avg.load_avg * cfs_rq->h_load, cfs_rq_load_avg(cfs_rq) + 1); So the load avg is -1, thus after adding 1 we get division by 0 The fix of the LKML reporter was to include the patches to kernel/sched/fair.c up to 4.5 A specific patch was not identified. Please backport these patches for Xenial and lts-xenial kernel in trusty. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/1568729/+subscriptions -- Mailing list: https://launchpad.net/~kernel-packages Post to : kernel-packages@lists.launchpad.net Unsubscribe : https://launchpad.net/~kernel-packages More help : https://help.launchpad.net/ListHelp