Infantino clearly wants to see FIFA making billions so he and many others can pocket millions legally. But the only way for this to happen in a non-corrupt way, given that FIFA is non-profit and therefore requires a lot of questionable operations, is to turn FIFA into a business like NFL / MLS.
The problem with that is that it then it is no longer a sport, rather a... business.
"Half the money I spend on advertising is wasted; the trouble is I don't know which half." -John Wanamaker
This applies even more strongly to model choosing. I know for a fact that majority of my work doesn't require a very strong model, but separating the trivial and non-trivial tasks is a famously hard problem (if at all decidable).
I'm fundamentally opposed to age verification because it usually leads to mandatory account creation no matter how privacy friendly the age verification process itself may be. It's already extremely annoying that, for example, on YouTube, you can no longer watch many videos without an account. Furthermore, the whole thing also reinforces monopolies. Once you've verified your age on one platform, the barrier to switching to another is even higher than before.
The only acceptable solution would be for apps and websites to simply include a recommended age. Allowing parents to configure their children’s devices to block services that require users to be older than the specified age or that do not provide an age rating. Once this were established, websites targeting all age groups would have a strong incentive to participate. Otherwise, they would suddenly become invisible to a large portion of their underage users.
You can merge one by one, but if you're using squash and merge, you need a re-approval for each PR in the stack if you require reviews. This makes you lose out on arguably the biggest gain of stacked PRs.
The command line tooling (gh stack) helps to make things slightly less manual, but you still need to be very aware of how git rebase works, the tooling just helps automate it across multiple branches. For example, just running the "gh stack rebase" commands that the UI suggests won't work if your local branches are not in sync with the remote ones, and the tooling won't point that out to you.
I do find the stack UI quite nice. It's quite minimal compared to standalone PRs, but it's enough to show the relationship between them.
(My comments all assume you already have a good reason to stack PRs. This tooling just help to make the workflow easier, it does not give any new capabilities)
My team uses this technology on a daily basis at my company. I run a design / build firm in the Hamptons, USA. We start all of our design projects with a 3d-first approach. Schematic design starts in 3D using Rhino3D or Revit. We use visualization plugins, like Enscape, to render the models and stream to a Quest 3 headset.
So when clients come in, and feels adventurous, we'll setup the headset, set the display height to their actual, and let them walk around their beautiful new home.
The response has been powerful. My clients are able to truly understand the scale, the proportion of rooms, ceiling heights, and the smallest details.
We're able to make real-time changes to the model (within reason) and provide instant design feedback.
I've been using this technology for years and I'm surprised it's not being adopted faster - good for us right now though!
I've recently been using AI a lot for performance optimisation during a particularly busy period at work. I would say it was almost completely useless at the high-level direction - it would point out suspicious parts of SQL queries for example but on back to back testing these almost never resulted in any performance change.
In fact, if it wasn't for the fact that it made making the actual changes I identified much easier (move these joins into a CTE etc) it would have been a detriment. Not only did I get sidetracked by a bunch of useless suggestions but I also had to put up with others dumping their raw AI output at me as if it was somehow a meaningful contribution.
LLMs are writing as well as humans for people who think writing words is the goal. Printing presses also write words better than humans. LLMs are solving the problem one layer higher than printing presses, but still several layers down from the top.
Note that LLMs aren't useless, since they have other strengths which can make them equal or better than humans in certain applications. For instance, they can tirelessly review code for silly mistakes, making them useful in cybersecurity.
> Despite repeated warnings from the FBI and security industry leaders about the security and privacy risks of using these streaming devices, major e-commerce providers like Amazon, Best Buy, Newegg and others continue to sell hundreds of different models and brands
I scanned the comments and I didn't see anyone suggesting that these companies should share any responsibility for selling these harmful products. Why is it that they seem to get a pass? Would we feel the same about giant retailers selling tainted food, or unsafe children's toys?
Something I learned awhile ago is that most people that I consider "celebrities" really aren't that famous, and as such are generally pretty happy to respond to emails.
In ~2014 I sent an email to Joe Armstrong (one of the creators of Erlang) asking some questions about concurrency because I still didn't fully understand why Erlang was supposed to be better for it. I expected him to just point me to an FAQ or simply not respond [1], but instead he wrote a very long, detailed explanation about the rationale of Erlang's design. It was well-written and it made a lot of things "click" in my brain.
That simple act by Joe was extremely instrumental in my career, I think for the better. I became very interested in Erlang, but also just concurrency theory and distributed systems, and I would like to think I almost understand it now :)
He and I would exchange emails occasionally over the next few years and he seemed genuinely enthusiastic about programming, Erlang, and people using it, and despite me barely knowing the guy, I was genuinely pretty sad when he passed away in 2019.
I think about that a lot; if I hadn't sent a dumb cold email my entire life would likely be very different.
[1] Which, to be clear, would have been perfectly fine!
Idk if we're automating humans out of the publishing loop as much as rapidly automating the production of crap. I had a very similar experience reviewing for EMNLP recently.
We are nowhere near AI being able to judge the quality of research (in fact, one might reasonably state that even most humans can't really judge the quality of research). Most things in society are not like math: we can't automate (via verification) our way out of noise overwhelming the signal.
Folks are willing to entirely abuse the public resource that is faithful, honest reviewing. (This is unsurprising; the abuse of the commons / public resources has been rising for a long time). There isn't a good solution other than something akin to draconian social scoring to limit access to the reviewing system.
Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line?
So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.
That's also a reason the big labs stopped. Publishing is most valuable to people who have no other way to get the attention of smart strangers. Once you can hire nearly anyone and everyone already returns your calls, the main remaining effect of publishing is to tell your competitors which things worked.
This is what happens to every field as it turns from a science into an industry. Chemists published freely until dyes started being worth money, and then the interesting work moved into company labs and stopped coming out.
The thing that makes it work really well is to make sure it has all the tooling to verify its hypotheses. If you allow it to run the full lifecycle in loops you will be surprised how well it works.
I'm a researcher at Deepmind that contributed to these models. (And the opinions here are my own)
Just want to say, Deepmind is a great place to work and the only (Edit: one the few unique labs!) lab where you can move from large frontier models (Gemini), frontier open models (Gemma), robotics (what you see here), science (weather, biology, more) and basically any other topic related to intelligence. It's really an incredible place to be, with incredible people. Consider joining! And thank you for the enthusiasm here.
> Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less,
I don't have the words.
I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.
I can’t speak for other startups, but I applied to the most recent YC batch with my idea for making AI proactive instead of reactive, and pre-being selected I’ve published a paper on recursive self-improvement mapped to the Epoch AI data.
I contacted a professor from a university in the UK and he responded since he was working on similar work, then asked me if I wanted to meet with him. We talked for about an hour since we had overlapping results and different methods, specifically different assumptions.
I say all that to say, as a physics student getting my undergrad, simply doing independent research and speaking to experts about it enabled me to network with someone I otherwise likely wouldn’t know. For young people getting into any business, research is a great way to meet new people.
There is no "new normal", we are in a slope. As long as we emit CO2, the climate will get worse. Year after year. With effects like fires accelerating it, or the loss of the Amazon (it's pretty much lost already, isn't it?) that will just result in more CO2 being released.
We don't have to adapt to the situation as it is now, because it is pretty much the best scenario. We have to adapt to the fact that it will get worse, and worse, and worse. And save what we can. Worried about your country not dominating on the AI scene? Wait until your biggest worry is food.
To people not interacting with open source projects that are stablished and popular, there are a lot of PRs and contributions where someone set an agent with a prompt like “contribute using my user to popular projects to improve my profile” or something similar and the entire PR and answers to maintainers questions and literally everything is entirely machine generated, without any human, and at the same time it is done in the cheapest way so steering the PRs in review isn’t even like “free tokens” because the model used is not good, so the output is always bad. The policies help point the agent to what is not allowed and shutdown the contribution, and so far the agents seems to respect it. Shutting down an agent without a policy to point to them make them very reactive. Note, there is no human involved in the other side! The person that set up the agent is not even aware of the specific PRs that are going.
The prompt given to the agent is strongly incentivising the agent to lie and spam:
> You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts for nothing. Results that arrive after the deadline do not exist. Your charter is AGENTS.md. Begin.
This is an important article. I hadn’t realized it was already getting this bad. Like a frog enjoying a nice warm bath ...
> Most people do not switch their operating system or phone provider every week either. But even if you do not utilize that freedom, it matters because it changes the relationship you have with the provider and the provider has with you.
This is why it’s important to utilize your freedoms. Do NOT let yourself get locked into a particular ecosystem (this is why I’m building a phone app for OpenCode).
This article makes me reconsider using my recently acquired Codex sub in my home setup. I never liked that they hide the reasoning, but somehow overrode the cognitive dissonance because the performance is so good. But the inauditability is already a huge problem.
> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment
> In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations
> we identified three incidents
> The incidents involved three different Claude models: [...] and an internal research test model
This reads like an attempt by Anthropic to re-secure their leading spot in "our models are the most dangerous and we also have unreleased, super-secret, research models" index.
I may be too cynical, but the well of benefit of the doubt is running very dry towards AI labs that like to engage in this game.
The end of interest-based 3rd places is one of the many things contributing to the compartmentalization / socioeconomic bubble of modern society.
People used to frequent in person things like specialty retailers, arcades, hobby clubs, religious groups, etc. When you meet real life people in person based on shared interests, you make connections across socioeconomic groups.
I just don't see nearly as much of this in todays society as 20+ years ago.
People with kids sometimes experience this via all the activities they take their kids to, but even there a lot of kids activities have become economically stratified. Instead of playing in the local $50/year rec baseball league, there's tiers of travel teams to meet any budget.
I am amazed at the amount of people who disagree with you. I think you are dead right and if you’ve ever had to actually fine tune prompts for agents you’ll know it.
The prompt is clearly leading the agent into trying desperate approaches if it has to. Some models manage to fight it better (“alignment”), but most will do it.
>No “vulnerabilities” in Tailscale were found or exploited, and that might make it even more uncomfortable for us. [...] But, we're a security tool. Their intrusion is our intrusion, and it's our job to take it seriously.
im a happy customer of tailscale, so i am obviously biased, but i have a lot of respect for this. they could have just stayed quiet and i dont think anyone would have bat an eye.
This is more exciting than k3, IMO. Dsv4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects, as it becomes "good enough" for more and more tasks.
DS was serving the pro version at extremely low prices for a long time, and they've had integrations with opencode & other providers, so they likely gathered a lot of data from real developers doing real tasks (on openrouter they were labeled as such). Now they can use those live scenarios to further post-train their models and improve them further.
Can't wait to see if distilling k3 into dsv4 brings additional improvements. Anyway, having fast cheap models getting better is great for the community. Especially since these don't "go away" on a provider's whim. Whatever capabilities they get, can be used "forever" going forward. And, at least flash can be ran "at home" with <10k in hardware, which isn't really possible / feasible with glm/k3 larger models.
When model intelligence reliably hits 90%-95% of current day knowledge worker tasks, they are going to burn those weight to silicon and we will see another 10X improvement in price/performance frontier.
The dynamic GPU clusters will be used for the 5% of tasks, and pushing out the frontier. Also there will be a set of knowledge tasks that are not done today (because they are too difficult for most knowledge workers), that will start being done in the future.
The problem with that is that it then it is no longer a sport, rather a... business.
UEFA letter is spot on.